Compare commits

...
206 Commits
Author SHA1 Message Date
dtourolle 9d04ff2154 Release 0.13.1
Benchmarks / CPU and I/O (per commit) (push) Successful in 12m24s
Benchmarks / Frame budget (on demand) (push) Skipped
🐳 Android image / Build and push (push) Successful in 4s
Build and test / android-image (push) Successful in 5s
Build and test / Desktop (Linux) (push) Failing after 31m25s
🐳 Windows image / Build and push (push) Successful in 4s
Build and test / windows-image (push) Successful in 4s
Build and test / Layer separation (push) Successful in 46s
Build and test / Android (aarch64) (push) Failing after 6h25m37s
Build and test / Windows (x86_64, cross) (push) Failing after 3m17s
Traceability / Requirement traces (push) Failing after 48s
2026-09-19 22:06:06 +02:00
dtourolle c6cfb2a02a Put the -1 on the greens along the chroma axis, not across it
The Malvar "R at green in R row" kernel weights the two greens two
sites away along the row at -1 and the pair up and down the column at
+1/2. The shader had the two swapped, in the comment as well as the
code, so the transcription checked against itself. Both sum to zero
and reconstruct a flat patch exactly, which is all the tests fed it.

On an edge the correction at green sites is half strength and the
false colour doubles: 0.375 against 0.19 on a grey step, and a
blue/yellow zipper around every clipped highlight at 1:1. The other
three kernels and the CFA tables were right.

A grey vertical step now runs through the pass; the transposed kernel
fails it at 0.375.
2026-09-19 22:05:39 +02:00
dtourolle b83f192847 Package release 2 of 0.13.0: the inference engine and the user runtime directory 2026-09-19 21:20:53 +02:00
dtourolle ecb648818b Search the user's own runtime directory before the system library
The reference desktop's only system ONNX Runtime is Arch's
onnxruntime-opt-cuda: 1.29, built without TensorRT and against cuDNN 8
on a cuDNN 9 machine. The probe rejects both providers correctly and
the app runs on the CPU provider, which is right and not what anyone
wants. runtime/ beside the models is now searched ahead of /usr/lib,
tools/fetch-desktop-runtime.sh fills it with the four libraries from
the current onnxruntime-gpu wheel (cuDNN 9, TensorRT 10), and the
About caption lists every rung that lost and why, not only the first.
Verified: the app selects TensorRT from that directory with no
environment variable set.
2026-09-19 21:15:55 +02:00
dtourolle 5fbf8944d7 Count the filler in the APK's bundled-model array
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m35s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 31m59s
Build and test / Layer separation (push) Successful in 40s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / windows-image (push) Successful in 2s
Traceability / Requirement traces (push) Failing after 57s
Build and test / Android (aarch64) (push) Failing after 54m6s
Build and test / Windows (x86_64, cross) (push) Failing after 1h4m55s
The unpack list gained migan-512.onnx without its length following;
nothing on the desktop compiles that crate, and the first Android build
of 0.13.0 stopped there.
2026-09-19 20:55:38 +02:00
dtourolle b502a8ef90 Release 0.13.0
Benchmarks / CPU and I/O (per commit) (push) Successful in 12m2s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Successful in 1h41m17s
Build and test / Layer separation (push) Successful in 50s
Traceability / Requirement traces (push) Failing after 1m6s
🐳 Android image / Build and push (push) Successful in 16m7s
Build and test / android-image (push) Successful in 16m9s
🐳 Windows image / Build and push (push) Successful in 6m11s
Build and test / windows-image (push) Successful in 6m13s
Build and test / Android (aarch64) (push) Failing after 40m26s
Build and test / Windows (x86_64, cross) (push) Failing after 1h4m8s
2026-09-19 20:43:27 +02:00
dtourolle 6fbdb1d06f Regenerate the traceability matrix and gesture book after the rebase 2026-09-19 20:41:44 +02:00
dtourolle d04b8f6044 FR-MRG-4: the border is cropped or filled, the fill experimental; panorama.md §13 records what was built and measured 2026-09-19 20:41:22 +02:00
dtourolle 8baa46ff49 Let the merge example fill, wait for engines and dump the filler's input; add a fill example that re-runs it stage by stage
A fill that went wrong took a seven-minute merge to look at again. Now
DR_FILL_DUMP=dir makes the merge write what the filler was given, and
the fill example runs fill_border on that, or a crop of it, on the engine
and writes coarse, each band and the feathered result as PPMs — seconds
per attempt on TensorRT. Both examples take DARKROOM_ORT_DIR as the app
does, and --wait-engines lets a compiling rung finish before timing.
2026-09-19 20:41:22 +02:00
dtourolle e43ae10439 Offer the border fill on the merge page, experimental, with every knob on it
A Border choice beside the projection — crop to the picture, or fill it
— that redraws the preview filled so the invented pixels are seen before
they are confirmed (FR-MRG-1), greyed with the reason when the model is
not there. The job fills at half the composite's resolution in a
display-ish space (white balance, matrix, gamma; invertible) and samples
the result back into the linear DNG wherever no frame reached; the
sidecar's merge line says border filled and with which knobs.

Experimental because the fill is right in thin borders and wrong in deep
corners, where the model's Places2 prior puts clouds in sky and water
under grass; so its six knobs — working scale, edge erosion, coarse pass,
band width, mirror depth, seam feather — are sliders under the choice,
each committing a redraw, until the defaults are right.
2026-09-19 20:41:22 +02:00
dtourolle 104e3a106f Fill a panorama's border with MI-GAN: mirrored context, coarse to fine, a feathered seam
dr_pano::fill owns everything the model does not — which tiles, what
context, how to blend — behind an Inpainter trait, and dr_pano::migan is
that trait over the shipped generator on the inference engine.

The known content is mirrored across the coverage edge into the hole and
a 256-px ring, the nearest 48 px folded, so the model interpolates between
real and mirrored sky rather than extrapolating into nothing. A coarse
pass at a quarter decides the structure with the whole border in a few
tiles; fine passes in 96-px bands from the edge outward texture it; the
seam is blended over a feather inside the real edge. Every knob is a
Params field, and an Observer hears each stage for whoever is looking at
why a fill went wrong.
2026-09-19 20:41:22 +02:00
dtourolle 031ba7b77d Hash a model's bytes once, at open, not on every acquire
A border fill acquires the filler once a tile, and each acquire hashed
the 28 MB model twice — 60 ms a tile, a third of the tile's run on a
throttled TensorRT. The Model keeps its hash from open.
2026-09-19 20:41:21 +02:00
dtourolle 54a80e688c Ship MI-GAN's bare 512 generator as the panorama border filler
Sargsyan et al., ICCV 2023; MIT code and weights (models/LICENCE.md),
exported by tools/export-migan.sh at a fixed 1×4×512×512 from the
authors' checkpoint — six operator types, 28 MB, in LFS like the rest.
The package installs it beside the scene model and the APK unpacks it
with the others.
2026-09-19 20:41:20 +02:00
dtourolle 2fd7690b6f Mark a pixel the lens correction pushed off the sensor with alpha 0 in the camera-space tap
The fused shader stored black with alpha 1 for a pixel whose source
coordinate left the frame, and the merge's warp averaged it in like any
other: a dark, badly interpolated fringe along every frame's edge, visible
as a seam wherever a frame ended and, later, as the edge the border fill
continued. The display keeps its opaque black; CameraLinear stores alpha 0
and the warp weights each sample by the alpha it interpolated, dropping a
sample that has none.
2026-09-19 20:41:20 +02:00
dtourolle c5f07f9ced Give the engine an Inpainter role for the panorama border filler
MI-GAN is plain convolutions, so every rung serves it and none needs a
special form; the role exists so resolve_model and the probe's fingerprint
know the model, and so the merge job can open it through the engine rather
than tract, which takes 7.4 s a tile for it.
2026-09-19 20:41:20 +02:00
dtourolle 5c00942b84 One completeness job over a registry of repairs, and a re-index button
A library's records are never all complete at once. A face found before
its quality was kept has no quality; one found before the eye models
existed has no reading; one adopted from a peer's shard has no crop; an
image the fast detector examined on a 1024 px proxy has boxes the current
detector would not have drawn; an image the scan stat'ed has no capture
date. On the reference library that is 17,762 faces under the bare
w600k_mbf id with no quality, no reading and no dense landmarks, 4,144 of
them without a crop, beside 12,217 images the fast detector examined and
found nothing in. Every one of those gaps was its own pass — V14's
measuring pass, §17.5's eye pass, the sweep's proxy repair, the sweep's
detector upgrade — with its own work list, its own count and its own idea
of done, and adding a per-face field meant adding a pass. There was no
pass at all for the case the library is actually in: boxes and landmarks
drawn by a weaker detector on a proxy, which every later per-face pass
would have read from.

dr_ui::repairs replaces them with one job over a registry. A Repair names
one thing a record can lack — the predicate that says which images still
owe it, the input its handler needs (a header, the original, or a native
render), the handler, and what to record for an image that can never be
done. The job unions the predicates into one work list, fetches each
image once at the most any claimant asks for, renders it at most once,
and runs every handler whose predicate that image still matches, checked
again before each because a detection writes every field a per-face
handler would fill. The registry today: face-proxy, face-quality,
face-eyes, face-crop, face-detection, face-upgrade, metadata — the last
there to say that this is not a face job. Adding a field is one entry.

A repair's predicate is the only definition of its work: the count the
settings page shows, the list the job fetches and the check before its
handler run are one predicate, so the job converges. That is why the
registry is cut to what the device can do rather than listing what it
skips — an entry is a count and a set of originals to fetch — and why an
eye reading that cannot be cut is not a criterion.

The catalog side is generic to match: record_updates writes whichever
fields a FaceUpdate carries and re-marks the image so the shards export
it; faces_needing and count_needing answer a predicate the caller
supplies, replacing the measuring pass's three special cases.

Two buttons on the settings page run the job and differ in one
predicate. "Index faces" converges on coverage: has anything examined
this image. "Re-index every face" converges on provenance: face-detection
claims every image with no marker under the chosen detector, in either
of its forms (FaceDetector::model_ids, so a desktop in f32 and a tablet
on the Hexagon do not re-index each other's work), and a marker saying a
weaker one looked is not that. An original over the fetch budget is left
exactly as it was under the re-index, where the sweep marks it examined:
a re-detection with nothing found would delete the faces, and "cannot
fetch" is not "no faces".
2026-09-19 18:52:13 +02:00
dtourolle 2a4ac0ed3d Carry every identity across a re-detection, by box and by embedding
record_detections replaces an image's faces and carried only the user's
confirmations onto the new ones, by box overlap above 0.5 IoU. Everything
else on the old faces was dropped: the suggestions the last grouping pass
made, and the people the user had said a face was not. On the reference
library that is 13,011 suggestions and 77 rejections beside 3,778
confirmations — a re-detection of it would have been correct by
FR-CULL-12's letter, since suggestions are derived data, and would have
handed back a People screen of strangers.

Now every old face is read before the delete — box, vector, assignment,
rejections — and matched to the new faces one-to-one, best pair first. A
pair qualifies when the boxes overlap at all and either the overlap alone
says so (IoU above 0.5, the old rule) or the embeddings do (cosine above
SAME_FACE_COSINE, 0.45, the reference library's P≈0.95 line). The
embedding route claims the box a low-resolution pass drew badly enough
that overlap alone would not; the vector is also what breaks the tie in a
group photograph, where two neighbouring faces overlap both new boxes.
Overlap is required on both routes, because the same vector elsewhere in
the frame — a mirror, a print on the wall — is not the same face and must
not take its name. Onto the matched face go the assignment as it was,
confirmed or suggested with its probability, and every rejection.

The merge's match_faces still matches by overlap alone across devices; it
is the same question and is not changed here.
2026-09-19 18:34:31 +02:00
dtourolle 46af2a0a46 Stop naming an optimisation level: on tract it means into_optimized, which aborts on yolo26n-seg
ONNX Runtime's default is already its fullest level. ort-tract maps any
level but disabled to tract's optimiser, whose slice pass divides by
zero inside the segmenter's graph — a panic across the C API and so an
abort, which is what stopped dr-ui's develop test. The app never asked
tract for that and does not start now.
2026-09-19 16:35:22 +02:00
dtourolle 95c9cffc0d Keep the embedder off the Hexagon, and let the probe example ask for a runtime
On the tablet the engine compiled arcface for the NPU: the routing
compared the form a rung wants with the form on offer, and for the
embedder both are f32, so nothing said no. A rung now says which roles
it serves at all, and the Hexagon does not serve the embedder (§7 —
its vectors must compare across devices). Tested at the routing seam.

dr-segment's onnx_probe example still named ort-tract, which is what
stopped the workspace test build.
2026-09-19 16:21:53 +02:00
dtourolle cbbe67fbd7 Let the probe's clock be its proof, not disable_cpu_ep_fallback
The strict flag refused the Hexagon over the ten quantise/dequantise
nodes at the graph's edges that QNN declines by policy, which cost
microseconds. A provider that hands real work to the CPU is slower than
the CPU floor and the timing already rejects it; the tablet measured
2.3 ms on the NPU against a 29.7 ms floor.
2026-09-19 16:13:19 +02:00
dtourolle 691af96e3e Keep the readable half of a provider's error for the settings row
ONNX Runtime's errors open with a source path and a template signature;
the first 160 characters of a CUDA failure were all signature. The
reason now starts at the first word a person can act on.
2026-09-19 16:07:19 +02:00
dtourolle 7a436e2549 Move the panorama keypoint detector onto the engine, and probe with a detector
XFeat's two exports are a Keypoints role now; the crate no longer names
tract, and the app compiles TensorRT engines for both ahead of the
first merge. The probe picks the smallest *detector* rather than the
smallest file: the tablet's first run chose the 112 KB eye classifier,
which has no int8 form, and reported the Hexagon as failed for want of
one.
2026-09-19 16:05:08 +02:00
dtourolle 76bc5652d7 Calibrate the int8 detectors on library proxies, in chunks, and measure them
The first int8 files found no faces at all, and for two reasons the
tool now guards against. The calibration set was landscape photographs
with no faces in them, so the score head's ranges had never seen the
face regime; the set is now proxies from the library itself. And ONNX
Runtime's strided and moving-average calibration modes both degrade
these graphs measurably (a quarter of the faces at eight images, none
at ninety-six), while driving the calibrator in chunks by hand gives
ranges identical to a single pass — so the tool does that, four images
at a time, and feeds quantize_static through its range cache.

Measured against f32 over 400 proxies (docs/inference.md §10.1): the
10g form finds every face above 32 px the f32 form finds; 500m and
2.5g find 96%, and what they lose sits at a median confidence of 0.52
against the 0.50 threshold. Shipped with the number on record.

The Android unpack list gains the three int8 files; without that the
tablet never saw them. D13's runtime half records the reopening.
2026-09-19 16:02:44 +02:00
dtourolle 4ed29b9d81 Add the int8 detectors for the Hexagon, calibrated on real photographs
tools/quantise-models.sh writes the QDQ form QNN's HTP backend takes
whole: opset 17, per-channel int8 weights, uint8 activations, ranges
from running the f32 graph over photographs fed exactly as the app
feeds them. The calibration is strided, four images at a time, because
every ONNX Runtime calibrator holds each image's whole set of
activations until it folds them — a gigabyte an image on the 10g
detector, and an OOM kill with no message when folded once at the end.

Release-time, never on the device (docs/inference.md §5): it needs
real photographs and a person reading the recall measurement that
gates whether each file is offered.
2026-09-19 16:02:37 +02:00
dtourolle 05508741af Start the inference engine from both apps and show its choice in Settings
The desktop names where a package may have put libonnxruntime — an
override variable, beside the executable, the package's own library
directory, the Flatpak prefix, the system library directory — and
Android points at the APK's native library directory, which is also
what Qualcomm's DSP loader must be told for the Hexagon skel. Android
starts the engine at the end of the model unpack rather than at launch,
because the probe fingerprints the model files and a first launch has
none until then.

The About panel gains an Inference row beside Graphics, re-read every
two seconds while the probe runs and engines land, and faces.model_id
carries the detector's form: an int8 detector finds a different set of
faces and is a different population (docs/inference.md §7). A
low-memory signal drops every idle session with the GPU caches.

The APK assembly bundles ONNX Runtime and the Qualcomm HTP libraries
from Maven, fetched by tools/fetch-android-runtime.sh with their
published checksums; RUNTIME_DIR=none builds the tract-only APK, which
is a slower app and not a broken one. The desktop packages carry no
runtime yet.

Two probe fixes from the first desktop run: the floor must not be
built with CPU fallback disabled, and a versioned libonnxruntime.so is
a runtime too. On the reference desktop the probe now loads ONNX
Runtime 1.30, measures 30 ms on the CPU provider, and selects TensorRT
at 1.5 ms.
2026-09-19 16:02:37 +02:00
dtourolle d15c41e699 Add dr-inference-engine and route every model session through it
One crate names the runtime, the providers and the devices; dr-face and
dr-segment ask it for a session by role. It hands ort an API table once
per process — from a libonnxruntime it dlopens when the app names a
directory holding one, otherwise from tract — so the Rust build stays
free of C on every target and a package can install the runtime as a
file (docs/inference.md §3).

Sessions live in a registry behind a Model handle that holds the bytes,
not the session: every use refreshes a timestamp and a reaper unloads
whatever sat idle past the decay. A scan that runs the detector on each
image never lets it go idle; a click in the develop view lets the
segmenter go after thirty seconds; a handle used after that reloads,
and reloads on a higher rung if a compiled engine has landed meanwhile.

The probe walks the platform's ladder by building strict sessions and
timing them against the CPU provider, caches the choice against a
fingerprint of the runtime, driver, hardware and models, and compiles
engines for the selected rung in the background, smallest model first.
Nothing in this commit turns the native path on: the apps still run on
tract until they call init with a runtime directory.
2026-09-19 16:02:37 +02:00
dtourolle caf21bea64 Name the crate dr-inference-engine 2026-09-19 16:02:37 +02:00
dtourolle 6739fdf908 Specify per-device inference backends, with the 2026-09-19 measurements
tract runs every model on one core on every platform. Measured against
ONNX Runtime's providers on the MagicPad 2 and the reference desktop:
ORT CPU alone is 3-10x, the Hexagon at int8 runs the detectors in
1-3 ms, TensorRT is ~2x the CUDA provider. NNAPI, XNNPACK, WebGPU and
CUDA int8 were tried and excluded with the numbers that excluded them.

The spec keeps the build C-free: ort::set_api takes a table from a
dlopened runtime or from ort-tract, chosen once per process. Rungs
are chosen by building a real session, cached until an input changes,
and compiled engines are built in the background after the first
frame. The embedder stays f32 everywhere; int8 detectors are a
distinct model_id and are gated on a recall measurement.
2026-09-19 16:02:37 +02:00
dtourolle 42d11d919b cargo fmt and clippy across the panorama work, and one lint master carried
The dr-face comparison is master's: a negated partial-order test on the
eye box's width, rewritten as the two conditions it meant.
2026-09-19 15:53:06 +02:00
dtourolle 67f225beba panorama.md: MI-GAN as the border filler — MIT, six operators, 7.4 s a tile
Read and measured, not built. The bare 512 generator exports at a fixed
shape and loads under tract with nothing unsupported; at f32 on the
desktop CPU it takes 7.4 s per 512×512 tile, which puts a full-resolution
fill of the fixture's border at ten minutes. The three routes that would
make it viable are recorded, with the quarter-resolution fill the cheapest
and Hexagon int8 the one the model was designed for.
2026-09-19 15:30:49 +02:00
dtourolle 39adfd4b75 Regenerate the traceability matrix and gesture book after the rebase 2026-09-19 15:24:44 +02:00
dtourolle 57ed51c1c5 Projection chips redraw the preview; auto-crop as the DNG default crop
Picking a chip stored the choice for the merge and changed nothing on
screen — the chip did not even highlight, since the selected property
was never written back. Now the pick is reflected, and the job, waiting
for its decision, takes a Preview request, draws the alignment on the
chosen surface at proxy cost and reports again; the drain puts the new
picture and its size up. Auto is the surface the field of view suggests.

Also:
The largest rectangle inside the frames' coverage is found a row at a
time — a histogram of consecutive covered rows and a stack pass per row —
so the composite is never held to be measured (FR-MRG-11). It is written
as DefaultCropOrigin/DefaultCropSize (FR-MRG-4): the file opens on the
picture, the border is still in it, and resetting the crop shows it.
rawler reports the crop as the picture, which the test checks.

FR-MRG-4 records the question raised the same day — fill the border
rather than crop it — as open: a non-generative fill through the heal,
or a generative inpainter with its licence and weights. Neither decided.
2026-09-19 15:24:20 +02:00
dtourolle 30bd276d0b Merge page: outline every frame on the preview, and let it be tall
A sweep whose frames overlap by more than half reads as one photograph,
and the page's job is to show frames. Each footprint is walked along its
border and drawn in amber where it lands, so twelve frames look like
twelve and a misplaced one is visible as such. The preview may take most
of the page's height rather than 320 px.
2026-09-19 15:24:20 +02:00
dtourolle 75d2ceb23c Provenance in the sidecar, a launch hook for the page, and where it stands
derived_from and merge are top-level sidecar fields (FR-MRG-6): one line
per source in order, and how the composite was made. A build that
predates them keeps the lines as unknown and writes them back. The job
writes the sidecar beside the composite and stages it with its own
record when the composite goes through the outbox.

DARKROOM_START_MERGE=a.CR2,b.CR2 lands on the merge page at startup with
the job running on local files, on the model of DARKROOM_START_IDENTITY,
for looking at the page where synthetic clicks do not reach it. The fetch
and the start are shared with the grid's button.

panorama.md §11 records what exists, the fixture's figures, and the six
things still open, auto-crop first.
2026-09-19 15:24:20 +02:00
dtourolle 2e9a1eb0f0 The merge job and its page: a selection to a panorama DNG, confirmed first
dr_ui::merge is the orchestration with no interface in it: decode each
frame to sensor data and build its graph as a session would (orientation,
lens profile); render each through the camera-space tap at proxy size and
detect keypoints there, so the alignment is measured in the undistorted
frame the tiles are rendered in; align; solve one gain per frame from the
proxies' overlaps; draw the aligned set in colour for the page; then wait.
Nothing is written until a Decision arrives (FR-MRG-1). The merge writes
a linear DNG through the outbox with a destination record, so the drain
puts it beside its sources on a folder library and a server alike, and
the library rescans (FR-MRG-3).

merge.slint is the page, on the import page's model: the alignment
table with a failed frame named on its row and the button held off
(FR-MRG-5), the preview, the projection choice, Stop and Back. A
"Merge to panorama" button joins the grid's selection bar at two frames.

Headless, the example produces the fixture's 22 993 x 5 980 DNG in 45 s
on the reference desktop, exposures balanced across the stop of drift.
2026-09-19 15:24:20 +02:00
dtourolle 44ea763c61 dr-gpu: the merge pass — warp, accumulate, resolve, chunk by chunk
merge.wgsl warps one camera-space tile into one output chunk — output
pixel to direction (the projection maths of dr_pano::projection, verbatim),
direction to the frame's camera, camera to source pixel, bilinear by hand
from four textureLoads because rgba32float is not filterable — and adds it
into a storage-buffer accumulator weighted by its distance from the
frame's edge. A resolve pass divides by the weights and packs sixteen-bit
samples at the sensor's scale with a coverage bit.

MergePass::merge drives it: bands of rows, chunks across a band, and for
each chunk only the frames whose footprint meets it, each rendered as the
source rectangle the chunk needs and nothing more. The working set is one
chunk, one tile and one band (FR-MRG-11); the frame textures are the
caller's to cache. Feathered, not seamed; gain a scalar per frame — the
blend quality is panorama.md §10's step 5, after the path writes a file.
2026-09-19 15:24:12 +02:00
dtourolle acab0d7abb A linear DNG in and out: the writer, and a three-sample RawImage
dr-export gains write_linear_dng — LinearRaw, DNG 1.4, u16 samples at
the sensor's scale, the body's matrices with their illuminants, the
as-shot neutral, the EXIF block an export writes — streamed strip by
strip through a closure so the composite is never held (FR-MRG-11). The
tiff crate's directory is a map, so PhotometricInterpretation is written
over what new_image set, which is the trick the S15.1 spike thought it
had to hand-roll around. The test reads the file back through rawler.

dr-decode's RawImage carries samples_per_pixel (a linear DNG is 3), the
body's profile with its calibrations mapped back to EXIF illuminant
codes, and the cleaned make and model. The GPU uploads a three-sample
image as it is, normalised by black and white like a photosite, through
a full f16 conversion — subnormals kept, because a 14-bit LSB sits at
f16's smallest normal and rounding it to zero would crush exactly the
shadows the file was written to keep.
2026-09-19 15:24:12 +02:00
dtourolle 9b6b4942cf The camera-space tap: OutputMode::CameraLinear, composed with no operations
compose_camera_linear composes the fused pass with an empty operation
list, the file's orientation as the baseline, a view rect for the tile,
and a store of rgba32float. On the GPU, render_camera_linear is the only
entry that accepts it: it fills the profile uniforms neutral — unit white
balance, identity matrix, curve off — so what lands in the texture is the
sensor's numbers after the lens warp and nothing else (FR-MRG-2). A third
bind-group layout carries the format, as the linear one does, and the
readback is generalised to any pixel width for the f32 copy.

Thirty-two bits because the composite is written back at the sensor's
scale: a 14-bit sensor has 16 384 steps to white and f16 keeps 2 048 of
them in the top octave.
2026-09-19 15:24:12 +02:00
dtourolle 54290b9540 dr-pano: a second XFeat shape for portrait frames, and a matcher that takes seconds
Twelve real frames from the fixture set now align in 4.5 s — 4.4 s of
matching, 118 ms of bundle adjustment — where the first run took 51 s and
left the first two frames out.

The matcher computes each pair's similarity matrix once, across the
cores, with a dot product written to vectorise; both nearest-neighbour
directions read it. The frames that failed were portrait: fitted into the
landscape input they used 512 of 1024 px, and their thin overlap did not
survive at half resolution. The same weights are now exported at 768×1024
as well and the detector picks the shape by aspect. The example aligns
from embedded previews and draws the set on a cylinder; on the fixture the
sweep is 152° at a fitted 47.9 mm against the EXIF's 50, RMS 1.5 px, and
the overlaps show no ghosting.
2026-09-19 15:24:12 +02:00
dtourolle 231b4a54ab dr-pano: the geometry, from features to cameras
A new crate holding the CPU half of a merge (FR-MRG-10): the grayscale
proxy with orientation, the XFeat decoder ported step for step from the
reference detectAndCompute, mutual-nearest-neighbour matching, a robust
pairwise homography with the focal length read off it, a hand-rolled
Levenberg–Marquardt bundle adjustment over every rotation and the focal,
the three output projections, and align(), which chains it all and names
the frames it could not place rather than guessing (FR-MRG-5).

Dependency-free without the xfeat feature — linalg.rs says why the dense
algebra is hand-rolled — and tested on synthetic sweeps whose answer is
known exactly. The noise test records the single-row degeneracy: one
pixel of noise is a tenth of a percent of focal, which is a uniform
stretch of the sweep, not a misalignment.
2026-09-19 15:24:12 +02:00
dtourolle 2bf0ec8dba S15.4, CPU half: XFeat runs in ~400 ms per frame on the tablet
tools/onnx-probe-on-device.sh cross-builds dr-segment's onnx_probe
without the embedded segmentation model, pushes it with a model to the
attached device and times two runs. The 768×1024 XFeat export takes
~400 ms on the reference tablet's NEON cores against ~300 ms on the
desktop, with identical output ranges — inside NFR-MRG-1's 1 s per frame.
The blend half of S15.4 waits for a chunked blend to exist.
2026-09-19 15:24:12 +02:00
dtourolle 5bf06c5030 Fixture README: the frames carry Orientation 8, not 6 2026-09-19 15:24:12 +02:00
dtourolle 44fdcbc6f7 Add the twelve-frame 6D panorama set as an LFS fixture
fixtures/pano/2025-08-05: _MG_8320 … 8331, one portrait hand-held sweep
at 50 mm with a stop of shutter drift and sky in every frame — the set
§3.11 is built against, with each of those facts named as the test it
is. fixtures/** is tracked in LFS like the models but with the opposite
default: CI's pulls exclude it, so a build never fetches 325 MB it does
not use.
2026-09-19 15:24:11 +02:00
dtourolle f9510405c3 FR-MRG-3: the composite is a RAW at the source's native scale
Camera-linear u16 samples on the first source's black-subtracted scale
with its white level, never rescaled to fill 16 bits, with its body,
matrices, illuminants and as-shot neutral carried — so the panorama is
developed afterwards as one photograph from the sensor's own numbers.
The only thing a warp cannot preserve is the colour filter array, and
the clause says so.
2026-09-19 15:24:10 +02:00
dtourolle 7e6b25b21b S15.3: the camera-space tap is uniforms, not structure — and FR-MRG-2 moves below the profile
The fused chain, as operation.rs's tests fix it, is warp → as-shot white
balance → operations → base curve → camera matrix → store. LinearWorking
stores after the matrix, so the existing linear tap carries the body's
base curve, and a composite stitched from it and developed as an
unprofiled body would render that curve twice.

FR-MRG-2 therefore stitches camera-linear RGB — after the warp, before
white balance, curve and matrix — and the composite carries the first
source's body, matrices and as-shot neutral so its own develop applies
the profile once. The composer already makes this a uniform question:
white balance, matrix and the curve flag are reserved uniforms, so the
tap is a compose entry with no operations and a render entry that fills
them neutral. panorama.md §5.1 states the shape and asks for f32 buffers.
2026-09-19 15:24:10 +02:00
dtourolle e4b6b6c935 S15.2: XFeat exports at a fixed shape and loads under tract
tools/export-xfeat.sh exports the convolutional network alone at 768×1024
grayscale, on the pattern of export-seg-model.sh: thirteen standard
operator types, no dynamic axes, the keypoint decoding left to Rust.
examples/onnx_probe loads it through the ort-over-tract backend the app
ships with nothing unsupported and runs it in ~300 ms on the desktop CPU.

The weights are Apache-2.0, read from the repository's LICENSE, with no
grant on the checkpoint — recorded in models/LICENCE.md before they land,
as FR-MRG-8 asks. The probe stays: the next model will need the same
check.
2026-09-19 15:24:10 +02:00
dtourolle 1ded5afbaa S15.1: rawler reads back a linear DNG, so that is the container
A hand-rolled 64×48 LinearRaw DNG — one IFD, 16-bit RGB, DNGVersion,
ColorMatrix1, AsShotNeutral — comes back through rawler 0.7 with cpp 3,
the samples in the order written and the matrix parsed into the camera
definition; CameraProfile::extract builds a profile from it. ImageMagick
reads the same bytes.

dr_decode::decode currently accepts the file as CFA and passes three
times the samples on, so the cpp == 3 branch is the decode work FR-MRG-3
needs, and the only decode work. panorama.md §8 records the result.
2026-09-19 15:24:10 +02:00
dtourolle c901fc1a0a Specify panorama merging: §3.11, D18, S15, and the design in panorama.md
A merge writes a new source file beside its sources (D18) rather than a
multi-source Version, which answers the schema question §7 had been holding
open for panorama, HDR merge and focus stacking together. The panorama is
undeferred as FR-MRG-1 … 11; the other two stay in §7 with their data model
decided.

FR-MRG-10 and 11 fix where the work runs — every per-pixel stage on the GPU,
the composite never held as one texture — because the output exceeds
max_texture_dimension_2d before it exceeds memory. panorama.md carries the
stage table, the chunked output driver, the model licences and the porting
sources. S15 gates all of it.

Coverage falls from 83.0% to 77.2%: thirteen requirements entered with no
code, and outstanding.md §11 says so.
2026-09-19 15:24:10 +02:00
dtourolle f79a76f2d5 Name the eye pass on the People screen
Once every image has been through the detector and only readings are
left — the state an already-indexed library is in the day the eye models
arrive — the button reads "Read eye state" rather than promising to
index, and the coverage line says what the faces are waiting for.
2026-09-19 14:24:16 +02:00
dtourolle facb44cb55 Keep the dense landmarks behind each eye reading, packed
The 106 points the eye boxes were cut from, stored beside the reading as
16-bit fixed point over the frame: 424 bytes a face, a seventh of a pixel
on a 6000-pixel frame, where f16 at the same size would have been six.
Derived data like the embedding, kept for the same reason — it cost a
fetch and a model run, and the next per-face pass should run from the
catalog. Shards carry it; a peer's shard from before it is still read.
2026-09-19 14:24:15 +02:00
dtourolle 85cc2b1dcc Trace the eye reading to FR-CULL-8a and the chip to FR-CULL-13
The register grew both clauses the same day this was built: FR-CULL-8a is
the per-face state the reading is, and FR-CULL-13 is the rule that a
signal is shown and filtered and never writes a judgement. The tags,
faces.md §17 and catalog.md now say which is which; FR-CULL-8a records
what of it is built, and that its third model is under the InsightFace
grant by the same decision as the pair.
2026-09-19 14:06:59 +02:00
dtourolle d706c12d77 Cover the eyes-open subquery with an index
The people filter was served from faces_image without touching a row;
reading the eye columns in the same subquery touched every one, and
ALTER TABLE had put those seven floats after the embedding and the crop
blob. One count took 24 seconds on the reference library, thirteen of
them system time. faces_eyes covers the subquery again: five
milliseconds.
2026-09-19 14:05:52 +02:00
dtourolle cd0ca6785f Specify eye state as a filter term, and record what was measured
FR-CULL-13, with §3.9.1's exclusion of blink detection re-read as the
exclusion of blink selection it always was: the stored fact and the chip
are built, a pass that picks the frame where everyone's eyes are open is
not. faces.md §17 has the models, the crop measurements, the four-state
rule and its floors, the native and proxy sheets read face by face, and
what remains to measure.
2026-09-19 14:05:50 +02:00
dtourolle 83f4253b6a Filter the grid to a person with their eyes open
An "Eyes open" chip beside the people chips, offered only while someone
is chosen and dropped when the last person goes, so no term narrows the
grid with nothing on the bar to say so. It compiles the rule in
dr_face::eyes into the person's face subquery — Anna, eyes open, whoever
else is blinking beside her — and drops a frame only on a closed eye that
could be read: sunglasses, eyes too small or soft to read, and faces never
read all pass, so an old library shows everything under the chip until
the measuring pass has run. A test drives the same readings through the
SQL and through the rule and requires them to agree.

The People screen badges a face "Eyes closed", "Sunglasses" or "Eyes
unclear" so the reason a frame is or is not in the grid can be read off
the face; the sweep loads the three models when they are beside the pair
and reads eyes on the indexing and measuring passes from the native
render; the coverage line counts unread faces as work to measure so an
already-indexed library keeps its Index button. The term travels with the
place.
2026-09-19 14:04:35 +02:00
dtourolle 6aae4c3eb0 Ship the three eye-state models beside the face pair
2d106det for the eye contours, OCEC for open or closed, SGC for
sunglasses — all three pinned to a batch of one by the same script as
the pair, and installed by every packager so the eyes-open filter works
out of the box. The two classifiers are MIT, code and weights; the
README records their provenance, SGC's undocumented training set, and
the hashes as fetched and as shipped.
2026-09-19 14:04:08 +02:00
dtourolle 54b543fb77 Store seven eye numbers per face rather than three
Per eye P(open), the pixels across its box and the sharpness of the
patch; and P(sunglasses). The verdict — open, closed, sunglasses,
unclear — stays a rule in dr_face::eyes so the floors can move without
re-measuring twenty thousand faces. Shards carry the same seven, and a
peer's shard from before any of them is still read.
2026-09-19 14:04:08 +02:00
dtourolle f5956707e7 Cut the eye box from a landmark contour, and refuse eyes that cannot be read
SCRFD's eye point places a face, not an eye: on turned and smiling heads
the classifier's window had the eye in a corner, and two model-free ways
of re-centring it — the darkest blob, the most contrasty window — both
lost open eyes (19 → 15 and 19 → 9 of 25). Three landmark models were
then run over the same faces; Face Mesh V2 and InsightFace's 2d106det
tied at 22 of 25 and 2d106det ships, being the cheapest by far and under
the grant the detector and embedder already carry. The eye box is the
tight bounding box of its ten lid points, cut upright from the native
render, which is what the classifier was trained on.

The larger change is that the reading now carries, per eye, the source
pixels across the box and the sharpness of the patch — because the
commonest wrong answer on the reference library was a soft eye read as
closed, and a classifier shown a smear will always say something. An eye
under either floor, or narrower than six tenths of its partner (the far
eye of a turned head, whose contour collapses), is not asked; a face with
no readable eye is a fourth state, Unreadable, that no filter drops. On
twenty native renders the one real blink is caught, the laughing faces
are closed, the profiles are judged on the near eye, and the one thing
left beyond any floor is a face with a pot held over it.
2026-09-19 14:04:08 +02:00
dtourolle b908d861e0 Keep each face's eye reading in the catalog and in its shard
Three nullable columns beside quality — P(open) for each eye and
P(sunglasses) — because the verdict is a rule with thresholds in it and a
rule belongs in code, not in rows that would have to be re-measured. NULL
is "never read": a face from before the models, or from a device without
them, and every reader treats it as unknown rather than as closed.

The measuring pass V14 built for the embedding's length is what fills
them, so the sweep's work list now also names faces with no eye reading
— but only on a device that has the models, or it would fetch every
original to do nothing to it. A peer's shard without the reading is still
adopted, unlike one without the quality: the pass finds this work by the
NULL rather than by the run marker, so adoption costs it nothing.
2026-09-19 14:04:06 +02:00
dtourolle 6b51726322 Read each face's eyes, and whether sunglasses hide them
Two MIT classifiers from the same author as the reference pipeline's
whole-body detector: OCEC answers P(open) for one 40×24 eye, SGC
P(sunglasses) for a 48×48 head. Both load in tract once their batch
dimension is pinned by tools/fix-face-model-shapes.sh, like the embedder.

The crops come through the same fitted similarity the aligned face does,
so an eye window is a constant in template units rather than a second
warp, and a tilted head yields an upright eye. Measured on 60 proxies
from the reference library: the eye window plateaus at 22×11, the S
variant beats M and L (which overfit their own domain), and for
sunglasses the aligned face beats a head framing but the higher of the
two catches 11 of 12 pairs against 9 for either alone.

The reading keeps both eyes and the sunglasses number apart, because a
wink averages to the least informative value and a lens of dark glass
draws a confident answer from the eye classifier — over a woman in
sunglasses it read the right eye 0.97 open. Sunglasses take precedence,
and a face behind them is neither open nor a blink.
2026-09-19 14:03:31 +02:00
dtourolle 2481904016 Bring outstanding.md up to the decisions of 2026-09-19
Its plugin section still asked for the contradiction to be resolved, its
render-path section still asked whether FR-DSP-2 was a requirement and
said NFR-RES-2 had no answer, and its closing section still called D12
open. Each now records what was decided and keeps the argument that was
weighed, so the document reads as the history it says it is rather than
as a plan the register has moved past.
2026-09-19 12:25:04 +02:00
dtourolle a92ae4576f Repair the references that point at sections that moved
Eight citations named §5.1, §5.2 and a §5 selector language that
requirements.md's §5 has not contained since it became a pointer at
architecture.md; two named §9 for the golden images and the benchmark
suite, which are §8; and the three pointers into architecture.md were
each one section off. All now name the section that holds the thing.

architecture.md §12's subsections are numbered 6.1–6.13, colliding with
its real §6. That numbering is what every ARCH §6.n citation in the tree
uses, so it stays, and a note at the head of §12 says so instead of
leaving the next reader to work it out.

FR-DEV-3f's open question about persisting the film stock was answered
in sidecar.rs; the clause now says so.
2026-09-19 12:25:04 +02:00
dtourolle ed4460cb9c Tag three requirements the code already meets
R5 says in its own note that zoom_resolution.rs establishes it as a
pixel equality; that file was tagged FR-DSP-5 alone. FR-DEV-19's three
sub-clauses carry eighty-three tags between them while the parent had
none; MaskLayer, which is the thing they edit, now carries it. And
NFR-R3 — a crash in decode does not take down the application, the
image is marked failed — is exactly what the decoder's panic guard and
the face sweep's unreadable mark do, tagged FR-RAW-4 and NFR-SEC-1 and
not the clause that asked for them.
2026-09-19 12:25:04 +02:00
dtourolle 7596cf9bcc State the compatibility baseline and the channels
NFR-COMPAT-1 and NFR-COMPAT-2 were instructions to write a requirement,
not requirements: "state the API level", "state the channels". Both
are now stated from what the build enforces and what exists.

The baseline is minSdk 28 / targetSdk 36 from the Android Dockerfile,
a Vulkan adapter at wgpu's default limits because compute needs storage
textures — device_from already called that the floor and is tagged for
it — with no optional feature required, since the f16 in FR-DEV-2 is a
texture format and not shader arithmetic. The reference device is the
HONOR ROD2-W09 the figures are taken on, and the second-vendor clause is
recorded as unmet rather than quietly dropped: there is no Mali or
PowerVR device, so an Android figure here is an Adreno figure.

The channels are all self-distribution — Arch package, local Flatpak,
sideloaded APK, NSIS installer — because D13's face weights rule out
every store, and the two consequences are written down: SAF stays
although a sideloaded build need not have it, and S11 becomes a
pre-publication step.
2026-09-19 12:25:04 +02:00
dtourolle 696bafa9d5 Undefer AI subject masking, which shipped, and give it a clause
§7 still listed "AI subject masking — deferred per D11" while
MaskSource::Subject and MaskSource::Category, backed by dr-segment's
instance and semantic models, had been the primary way a local
adjustment is made for weeks. The code was tagged FR-DEV-3, which
names gradients and brushes and says nothing about a model.

FR-DEV-3i now states what exists: a subject or a category found by a
local model, stored as identity with the run's signature so that it
merges per field and reads as stale rather than wrong, then treated as
any other layer by the edge, stroke, composition and reveal clauses.
The one place it departs from FR-DEV-19 — coverage written run-length
coded beside the layer, so a stored subject renders without a model —
is recorded in the clause instead of left for the next audit to find.
The segmentation crate and the UI's selection module are tagged to it.
2026-09-19 12:25:03 +02:00
dtourolle d259c0d4bb Say that FR-DSP-2 is waiting on S6, not that it was rewritten
R5's note said FR-DSP-2 "was rewritten rather than implemented". It was
not: the clause still demanded viewport tiling, the matrix listed it
unbuilt, and frame-budget.md's rewrite had been proposed and never
applied. Decided 2026-09-19 to keep it as written until S6 runs on a
mid-range Android device, because the measurement that argues against
tiling was taken on a discrete desktop GPU and the clause exists for the
device whose memory the image exceeds. Both notes now say that.
2026-09-19 12:25:03 +02:00
dtourolle 95458356da Record D13's position on the face weights
The licensing half of D13 had been open since 2026-08-09, while the
InsightFace detectors and embedder shipped in the tree and indexed real
libraries. models/face/README.md already stated the position the
project was actually taking; the register did not.

Now it does: this is non-commercial software, self-installed, and it
uses the weights under their research grant as such. The risks are
written where the decision is — the grant binds every user, it is not
GPL-compatible, it rules out every public channel, and publishing is
what reopens the decision. S14's licence search is what would close it.
2026-09-19 12:25:03 +02:00
dtourolle dc9db11033 Decide NFR-R8: no CPU pipeline, a degraded mode instead
NFR-R8 carried the words "decide explicitly" for six weeks, asking
whether v1 has a full CPU render path or whether "CPU fallback" means
staging only. The viewer had already answered it: with no adapter it
opens the library on embedded previews and cached proxies, keeps every
catalog edit available, and withholds develop and export. That is the
degraded mode, it is now the requirement, and NFR-RES-2 no longer
promises a fallback render path that was never going to be built.
2026-09-19 12:25:03 +02:00
dtourolle 3692306fd3 Say once that there is no phone
Three statements disagreed. §1.2's platform table said "phone
supported"; §3.5 said phones were out of scope and cited §1.3, which
does not mention them; D15 said "no phone" and gave the reason. D15 is
the decision, so the other two now point at it and say the same thing:
the build runs on a phone, and nothing is designed for one.
2026-09-19 12:25:03 +02:00
dtourolle c826fed605 Put the plugin API post-v1, and let the matrix count it that way
The register said two things about plugins. §7 had listed "Plugin API"
as deferred since the first draft, in a bare row; §3.10 then specified
it in 23 clauses that counted against coverage. Twenty-one of them had
no implementation of any kind, and could not have: no crate loads
anything at runtime. The coverage figure was measuring the contradiction.

Decided 2026-09-19: §7 is right. §3.10 stays as the design of record,
each of its clauses is marked "(post-v1)" on its defining line, and
NFR-SEC-6 — which exists only for plugins — goes with them, as does D16.

The traceability tool learns the marker. A deferred requirement is still
defined, so a tag naming it is not an orphan, but it leaves the
denominator and is listed in its own table rather than under "not yet
tagged". The marker must sit on the definition line; a mention of
"post-v1" in prose changes nothing, and where an ID is defined twice the
deferral on either line wins. Both are tested. Coverage moves from 72.2%
of 194 to 80.6% of 170 without a line of application code changing,
which is the honest figure: it now measures what v1 owes.
2026-09-19 12:25:03 +02:00
dtourolle c921852d89 Record D12 as settled by events, and D3 as delivered
D12 had been OPEN since the 2026-08-08 calibration, and D3 said it
depended on D12. In the meantime the milestone D3 named was delivered
and closed on 2026-08-30 and the application reached 0.12.2 with every
cluster the calibration selected at least begun. The decision the
register was waiting for had been made by building, so the register
now says so: full scope stands, v1 has no date, and "post-v1" in §7 is
the one way a clause leaves the count.

The status line also stops calling this a draft from August; it has
carried eleven dated amendments since.
2026-09-19 12:24:54 +02:00
dtourolle e6ac31d39d Give the face sweep a size budget, so a panorama is never fetched
The sweep fetches the whole original before it can learn anything
about it, and the one file in the reference library the decoder
refuses on sight is a 521 MB stitched panorama — so every pass on the
tablet spent half a gigabyte of Wi-Fi to find that out again. The
catalog already knows the byte count, and that is enough to decide
before the fetch: originals over 256 MB are marked examined with
nothing found and a zero edge, counted as failed, and named in the
log. Below the line is every camera RAW the library holds; above it,
four files, all panoramas.

A budget and not a verdict on panoramas. The right treatment for one
is a tiled pass — read it in strips, detect in each, stitch the boxes
back — and the zero edge is what that pass would select on. Until it
exists, this is what keeps a background sweep on a phone from paying
for the decision the decoder cannot make.
2026-09-19 12:14:05 +02:00
dtourolle 7db999c1f6 Require judgement anywhere, and evidence that never becomes a verdict
Rating and flagging were reachable from the grid alone, so a photograph
opened in develop could not be judged without leaving it; FR-UI-5 said
"rating" without qualifying the view and was built as though it had.
And FR-UI-1's expanded row has said "filmstrip" since it was written
while the roll stayed on demand in both classes. Both are amended to
say what they meant: judgement follows the photograph, without
auto-advance outside the culling mode, and the roll is open by default
where there is room for it.

The larger change is a rule. Per-face signals — eye state from a
classifier, head pose from the five landmarks the detector already
yields — are worth having for culling, and §3.9.1 excluded detecting a
blink outright. The exclusion was always of judgement, not of knowing:
a blink is a fact about a frame of the same kind as a clipped
highlight. FR-CULL-8a specifies the two signals; FR-CULL-13 says what
any signal may do (be shown, filtered, sorted, propose a burst
representative) and what none may (write a rating or flag without a
user action between). R7 states the same thing as a user need.

Licensing was read before either was written. OCEC's eye-state weights
are MIT with a clean data chain; every open gaze model is trained on
Gaze360 or its peers, whose licences restrict derived models by name,
so gaze is deferred in §7 and head pose stands in for it. D13 records
both so they are not re-searched.

Replacing a closed-eyed face from a neighbouring frame was raised and
is written down as D17 rather than built: it is the multi-source schema
question §7 already defers for panorama and HDR, with its non-goals —
never automatic, provenance declared — fixed now.

Traceability regenerated: three new IDs, none yet tagged.
2026-09-19 11:47:27 +02:00
dtourolle 30b89ad70a Merge: one face population per embedder, whichever detector found them 2026-09-19 10:52:31 +02:00
dtourolle 327decfab1 Fuse every detector's faces into one population per embedder
Choosing "Thorough" made the library look empty. The detector setting
writes under its own faces.model_id, and every reader of "the faces"
keyed on that exact id: the clustering pass, the coverage figure, the
sweep's work list, the shard export and import, and the sync merge's
face matching. On the reference library that restarted coverage at
1,834 of 19,140, drew a People rail of 36 faces for a person with 520,
queued a ~400 GB re-fetch on each device, and stranded the desktop's
3,583 confirmations under the old id: the tablet held the same faces
under the new one and the merge refused to match them. Same photograph,
same box, same embedder, two ids — that is one face, not two libraries.

The embedder half of the id is now the key. embedder_of and embedder_sql
give it to every query; writes keep the full id, so which detector drew
a box stays on record. record_detections is unchanged and is where the
generations meet: an image holds one pipeline's faces at a time, and a
re-detection carries confirmations across by box overlap. The merge's
match_faces applies the same rule within an embedder. The calibration
is keyed on the embedder too, since the similarity space did not change.

Shards travel every generation, each under its own id, and a peer adopts
whichever it is sent — including a stronger detector's pass over an
image it indexed itself with a weaker one, which is the re-detection its
own sweep would otherwise queue, already done. Never downwards: a tablet
on Fast keeps the desktop's Thorough faces. The sweep gains the same
tail — images a weaker detector indexed, after the ones nothing has —
driven by FaceDetector::supersedes, so choosing a stronger detector still
improves the library over time without first making it disappear.
2026-09-19 10:49:28 +02:00
dtourolle f8addbee53 Mark a file the decoder cannot open, so the sweep stops fetching it
A decode failure in the face sweep was counted, logged at debug where
nobody saw it, and left unmarked — so the next pass fetched the same
file and failed the same way. For the 521 MB panorama behind rawler's
panic that was half a gigabyte per sweep, on a tablet. It is now marked
examined with nothing found and a zero edge, which is what a later "try
again with a better decoder" pass would select on, and the warning
names the file. The failure count is unchanged: it did fail.
2026-09-19 10:45:08 +02:00
dtourolle c0b1e78f7c Return a panic inside the decoder as an error, not as the end of the thread
rawler panics on some input rather than returning Err — a DNG whose IFD
claims a >50000 px image, which the reference library has: a 521 MB
stitched panorama, IMG_4181-Pano.dng. On a worker thread a panic is the
end of the thread, so the face sweep that met it stopped thirteen
seconds in, three sweeps running on the tablet and three on the
desktop, with "17301 image(s) to index" as the last word. FR-RAW-4
says a malformed file must not abort a batch, and that is this crate's
promise whatever the library beneath it does: every entry point that
calls into rawler now runs under catch_unwind, and a file that panics
the decoder is one failed file with the panic's message in the error.

Verified on the panorama itself: metadata reads, decode returns the
error, the thread survives. The crash hook still records the panic,
which is right — it is a defect in a dependency and the record is how
it gets reported.
2026-09-19 10:45:07 +02:00
dtourolle 78cb00634e Fetch the photographs around the open one ahead of the step to them
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m59s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / android-image (push) Canceled after 0s
🐳 Android image / Build and push (push) Canceled after 0s
Build and test / Android (aarch64) (push) Canceled after 0s
Build and test / windows-image (push) Canceled after 0s
🐳 Windows image / Build and push (push) Canceled after 0s
Build and test / Windows (x86_64, cross) (push) Canceled after 0s
Build and test / Layer separation (push) Canceled after 0s
Build and test / Desktop (Linux) (push) Canceled after 16m26s
Traceability / Requirement traces (push) Canceled after 0s
Walking the photo roll was one download per frame: every step showed
"Downloading…" over an empty canvas while tens of megabytes came down,
and moving between a pair of near-identical frames paid that a dozen
times. Now, once the opened photograph has landed, the ones around it
are fetched into the originals cache while it is being looked at, so
the next step is a disk read.

A single worker serves the latest wish only, closest first and working
outwards — next, previous, next-but-one, previous-but-one… — one file
at a time. Each open replaces the wish, so a fast walk never leaves a
trail of stale downloads competing with the one being waited on. A
process-wide in-flight registry makes a click on a photograph that is
still being fetched ahead wait for that transfer and read it from disk,
rather than start a second download of the same file.

How far each side is a setting under STORAGE — Off, 2, 5, 10 or 20,
defaulting to 5 — and it is moot while "keep originals after opening"
is off, since a fetch the cache would discard on arrival is transfer
for nothing. Nothing is fetched ahead while offline. The transfers show
in the activity list while they run and are removed when they end.
2026-09-19 10:36:50 +02:00
dtourolle 2917b7427d Release 0.12.2
Benchmarks / CPU and I/O (per commit) (push) Successful in 12m51s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Successful in 1h44m42s
Build and test / Layer separation (push) Successful in 1m4s
🐳 Android image / Build and push (push) Successful in 6s
Build and test / android-image (push) Successful in 7s
🐳 Windows image / Build and push (push) Successful in 3s
Build and test / windows-image (push) Successful in 3s
Traceability / Requirement traces (push) Successful in 2m15s
Build and test / Android (aarch64) (push) Successful in 1h4m7s
Build and test / Windows (x86_64, cross) (push) Successful in 1h12m17s
2026-09-19 10:15:46 +02:00
dtourolle 33e2e277a2 Set the Wayland app id late enough for it to take
The launcher and the task bar have shown a generic tile for a working
window since the call was written. set_xdg_app_id sat at the top of
run(), on the reasoning that the app id is read when the surface is
created — true, and beside the point: the call goes through Slint's
global context, and there is no global context until something installs
a platform. That is BackendSelector inside shared_gpu, or AppWindow::new
falling back to the default, and both happen further down. Called before
either, it returned NoPlatform and did nothing at all.

It moves to just after the window is constructed, which is not the same
as shown — run() is far below — so there is a platform to talk to and
the surface does not exist yet.

The failure was logged at debug, which is why a year of grey squares went
unremarked: the whole symptom is invisible from inside the application.
It is a warning now, naming the consequence.
2026-09-19 10:15:46 +02:00
dtourolle 9b627e7713 Let a catalog writer wait for its turn instead of losing its work
SQLite's busy timeout defaults to zero, and nothing ever set one: the
loser of a write race got SQLITE_BUSY at the moment it asked. WAL does
not cover this — it makes one writer and many readers free, and this
application constantly has two writers, the face sweep committing a
batch while the derived sync imports shards or reclustering reads.

The cost was not a retry but lost work. A sweep that had already paid
for the detection and the embedding — seconds per image, the expensive
part — discarded the result on "storing faces for 214: database is
locked" and moved on to the next image. Both the desktop and the tablet
logged runs of those on consecutive images, which is a face sweep
quietly failing to store the faces it had just computed.

Ten seconds, on every connection, set in configure() so that nothing
can open the catalog without it — the figure the job runner's own tests
have used for this reason since they were written. It is far longer
than any transaction here, so it bounds pathology rather than making
anyone wait.
2026-09-14 20:05:53 +02:00
dtourolle 8c3b62745a Give makensis absolute paths, and one installer to find
Benchmarks / CPU and I/O (per commit) (push) Successful in 2m53s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Successful in 1h34m55s
Build and test / Layer separation (push) Successful in 1m10s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
🐳 Windows image / Build and push (push) Successful in 6s
Build and test / windows-image (push) Successful in 6s
Traceability / Requirement traces (push) Successful in 38s
Build and test / Android (aarch64) (push) Successful in 27m51s
Build and test / Windows (x86_64, cross) (push) Successful in 1h10m51s
The first CI run of the Windows leg passed every step up to packaging
and died in makensis with LicenseData: open failed
"target-windows/installer/stage\LICENSE". CI sets CARGO_TARGET_DIR to
the relative target-windows, and NSIS on a POSIX host translates the
backslash in a File path only when a leading / tells it the path is a
POSIX one; a relative name reaches it with the backslash intact.
Locally the target was always /work/…, which is why it never showed.
package.sh now resolves its directories with realpath first.

It also removes any installer already in the output directory before
writing the new one. That directory is cached between runs, so after a
version bump a glob over it finds two, and the smoke test hands Wine
both names joined by a newline as one path — which is what happened
locally the moment the version moved to 0.12.1.
2026-09-14 01:12:33 +02:00
dtourolle 8012979a1e Release 0.12.1
Benchmarks / CPU and I/O (per commit) (push) Successful in 10m58s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Successful in 1h40m20s
Build and test / Layer separation (push) Successful in 1m5s
🐳 Android image / Build and push (push) Successful in 4s
Build and test / android-image (push) Successful in 5s
🐳 Windows image / Build and push (push) Successful in 3s
Build and test / windows-image (push) Successful in 4s
Traceability / Requirement traces (push) Successful in 1m18s
Build and test / Android (aarch64) (push) Successful in 1h0m2s
Build and test / Windows (x86_64, cross) (push) Failing after 56m14s
2026-09-13 20:10:03 +02:00
dtourolle 5e67f026ec Merge: a catalog the server cannot damage for long, and backups on both sides
fix/corrupt-remote-catalog-deadlock. The catalog on the server had been
malformed since 7 September and every device declined to overwrite it,
so collections and people stopped syncing everywhere at once; the
damage came from two devices assembling chunks in one upload directory.
Each chunked upload now has its own directory, the server keeps three
generations of the catalog behind the current one, every push is
verified before and after, a damaged copy that arrived whole is merged
from the generation before it rather than pinned in place, and the
catalog is backed up daily as NFR-R2 always asked.
2026-09-13 20:04:55 +02:00
dtourolle eeee3d920a Back the catalog up daily, not only before migrations
NFR-R2 asks for the catalog to be backed up on a schedule and before
schema migrations. Only the second half existed: every backup on disk
was a pre-migration copy, and a library that never migrated was never
backed up at all.

A backup is now also taken at the end of a library sweep when the newest
one is more than a day old — the moment the catalog is quiet and a day's
collection and people edits have just been folded in — on its own
thread and its own connection, so the copy of a 130 MB file is not spent
on the UI. Whether one is due is read from the backup directory, not
the catalog, so the ordinary case costs nothing. An empty catalog is
skipped: there is nothing in it a rescan would not rebuild. Pruning to
KEEP_BACKUPS applies as before.
2026-09-13 19:32:05 +02:00
dtourolle f100db89ca Verify the catalog snapshot before it is sent, and after it lands
Two checks around the upload, both cheap next to what they prevent.

Before: the snapshot is quick_checked before it leaves. It is the copy
every other device merges from, and a damaged one costs each of them a
download, a failed merge and a refusal to push.

After: the staged upload's size on the server is compared to the bytes
sent before it is rotated into place. A chunked upload is assembled
server-side, and an assembly that goes wrong is a file of plausible
size no device can open — caught here, on the device that caused it,
for one listing; otherwise on every other device, after the fact. A
mismatch, or a size the server will not confirm, discards the upload
and leaves the current copy and its generations untouched.
2026-09-13 19:31:58 +02:00
dtourolle 2ce0fcc74a Keep three generations of the catalog on the server
The server held one copy of the catalog, overwritten in place on every
push. When that copy was damaged there were two answers, both bad:
refuse to touch it for ever, which is what every device did for a week,
or overwrite it with ours, which loses whatever another device had
added since — the escape hatch of the previous commit.

A push now uploads to catalog.upload.sqlite, rotates catalog.sqlite to
.1 (and .1 to .2, .2 to .3, dropping the oldest), and moves the upload
into place. Rotation is server-side renames, oldest first so that every
destination is empty when it is written to — move_to refuses to
overwrite, by design — and a failure at any step leaves a gap in the
generations and never a missing current copy. The only transfer is the
upload itself.

A damaged current copy that arrived whole now merges from the newest
readable generation before ours goes over it, which loses nothing, and
is kept as .1 by the ordinary rotation rather than by a separate 40 MB
upload. NFR-R2 asks for the catalog to be backed up; this is the half
of it that lives with the copy other devices read.
2026-09-13 19:31:58 +02:00
dtourolle 6f62ac09f8 Give every chunked upload its own directory on the server
The upload directory was named from the destination path alone, on the
reasoning that two files could then not collide. Two devices uploading
the same file could, and did: both wrote 00001…00009 into one directory,
and whichever MOVEd first assembled a mix of the two — a catalog of
exactly the right size whose pages came from two different databases.
SQLite called it malformed, every client declined to overwrite it, and
collections stopped syncing on all of them for a week. A transfer that
died on a phone's link left its chunks there for the next device to
assemble in, by the same mechanism.

The name now carries a nonce as well, so no two uploads share a
directory, and a failed transfer deletes its own directory on the way
out rather than leaving 5 MB chunks for the server to sweep eventually.
2026-09-13 19:31:58 +02:00
dtourolle c1e0f09be7 Say where the face models were looked for when they are not found
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m27s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Successful in 1h35m10s
Build and test / Layer separation (push) Successful in 39s
🐳 Android image / Build and push (push) Successful in 4s
Build and test / android-image (push) Successful in 4s
🐳 Windows image / Build and push (push) Successful in 7m11s
Build and test / windows-image (push) Successful in 7m12s
Traceability / Requirement traces (push) Successful in 58s
Build and test / Android (aarch64) (push) Successful in 24m57s
Build and test / Windows (x86_64, cross) (push) Failing after 57m30s
"The chosen detector is not installed" was the whole of what a user
saw, on a machine where the files were three directories away from
where the lookup went. The search order was in a doc comment and
nowhere a user could read it. Now a missing pair logs the detector
file it wanted and every directory it tried, which is what the first
Windows install needed and what the next misplaced download will.
2026-09-13 19:20:30 +02:00
dtourolle 38a87c0bca Put settings.json where the Android entry point said, not under /
SettingsStore::open read XDG_CONFIG_HOME and HOME itself. Neither
exists on Android, so it resolved to .config/darkroom relative to a
working directory of /, and every settings edit on the tablet failed
with "read-only file system" — the page reported the error and nothing
said why. The doc comment claimed "the same resolution SessionStore
does"; now it is, by calling the same function, which honours the
directory android_main declares and takes the platform's config
directory everywhere else. settings.json sits beside sessions.json on
every platform, as the comment always said it did.
2026-09-13 19:05:51 +02:00
dtourolle 693195fa96 Replace a damaged catalog on the server instead of pinning it there
The catalog sync refuses to upload when it cannot read the server's copy,
because the upload is a read-modify-write and writing blind would discard
another device's collections. That is the right rule for a timeout, a
dropped connection or a newer schema — the remote is fine, only our view
of it failed.

A file SQLite calls malformed is not that. No device will ever read it
again, so refusing to write over it preserves nothing — and every client
declines in turn, pinning the damaged file in place for good. Collections
and people then stop crossing between devices on all of them at once,
each logging "catalog not pushed" on every pass. This library did exactly
that from 2026-09-07, on the desktop and on a freshly installed phone
alike, while 32 collections sat undelivered.

Now a copy that arrived whole and still will not open is set aside under
a dated name and replaced by ours. Whole is checked against the size the
server advertises: a truncated download will not open either, and on a
phone that is the far likelier story, so anything short — or any size the
listing cannot confirm — is treated as the transport failure it is and
the server's copy is left alone. A placeholder's size is not trusted for
the comparison, since it means nothing.

The report says when this happened, and the log line calls it "pushed
over a damaged copy" rather than folding it into an ordinary push: it is
the one push that discarded something.
2026-09-13 19:01:00 +02:00
dtourolle 43f70c4765 Build the Windows installer in CI
The fourth leg of build-and-test.yml, in the shape of the Android one:
an image workflow that builds docker/windows and pushes it tagged by
the directory's tree id, and a job inside that image that lints the
Windows target — the only place the cfg(windows) branches are ever
compiled by CI — builds, runs the smoke tests docs/windows.md §6
specifies, packages, installs and uninstalls under Wine, and uploads
the installer. Every step was run by hand in the same container first.

The spec's open list closes with this: the four §3.2 items, the
licence page, and the leg. What remains is what Wine cannot show, and
§10 now lists it as the first real Windows run's checklist.
2026-09-12 07:34:10 +02:00
dtourolle 6609aa9acf Ship the GPL text, and show it in the installer
The repository declared GPL-3.0-or-later and carried no copy of it;
the Arch package pointed at the system's shared text and nothing else
needed one. The installer does: a licence page needs a file to show,
and the moment before installation is where the terms can still change
a decision. The standard text, at the root where every convention
looks for it, converted to CRLF at packaging time because a Windows
edit control draws a bare LF as nothing.
2026-09-12 07:34:10 +02:00
dtourolle b396096787 Open the sign-in URL on Windows
Login Flow v2 cannot complete without a browser, and the launcher had
a branch for xdg-open, one for Android's Intent, and an honest
Unsupported error for everything else — which on Windows stranded the
flow on "approve the sign-in in your browser". rundll32
url.dll,FileProtocolHandler is ShellExecute on the URL and needs no
crate; chosen over cmd /C start, whose quoting of & in a query string
is a known trap. Not verified: Wine has no browser to open.
2026-09-12 07:34:09 +02:00
dtourolle beb822dced Keep secrets in Credential Manager on Windows
The Secret Service store was keyring::Entry all the way down, and
keyring 4's v1 feature set — the one the workspace already asks for —
includes the Windows Credential Manager backend. So the Windows store
is the same implementation with its cfg widened, and the crate as a
target dependency. The one behavioural difference is that the
availability probe always succeeds there, which is correct: Credential
Manager is always present, so FR-NC-2's degraded mode does not arise.
Until now a Windows build compiled, started, and failed at sign-in
with the placeholder store's "no secret store is implemented".
2026-09-12 07:34:08 +02:00
dtourolle ef1154af94 Resolve every base directory in one place, and on Windows
Five sites each read XDG_*_HOME and fell back to $HOME/.local/… on
their own, which is fine on Linux and wrong everywhere else: Windows
sets neither variable, so every one of them degraded to a path
relative to the working directory — for a Start Menu launch,
C:\Windows\System32. The models lookup walked XDG_DATA_DIRS the same
way.

dr_plat::dirs now holds the rule per platform: XDG on Unix, the known
folders on Windows — %APPDATA% for config, which roams, and
%LOCALAPPDATA% for data and state, which do not — and the executable's
own directory as the system data dir, which is where the installer
puts the models. The Android overrides stay where they were; only the
fallback behind them moved. Both rule sets are unit-tested on either
host, and the Windows one was confirmed by running the application
under Wine: its log landed in AppData\Local\darkroom\state and nothing
was written anywhere else.
2026-09-12 07:34:05 +02:00
dtourolle 896188a489 Read the sidecars other editors write, and write them back on request
Benchmarks / CPU and I/O (per commit) (push) Successful in 10m59s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Successful in 1h33m36s
Build and test / Layer separation (push) Successful in 1m2s
Traceability / Requirement traces (push) Successful in 1m25s
🐳 Android image / Build and push (push) Successful in 9s
Build and test / android-image (push) Successful in 9s
Build and test / Android (aarch64) (push) Successful in 56m59s
FR-CAT-13 asked for standard XMP and `core/dr-xmp` answered the file: it
has read and written `dc:subject`, `xmp:Rating`, `xmp:Label` and the IPTC
core since 5fa4c07, under an ownership rule that leaves everything else in
the document untouched. What nothing did was call it. No scan found an
`.xmp` beside a raw, no catalog row was filled from one, no judgement
wrote one back, and the "external modification detected, reload offered"
clause had no mechanism. A library imported from Lightroom came in and
could not go back out.

The scan collects `.xmp` beside `.drsc` from the listings it was already
paying for, and the pull reads each one whose ETag has moved. Both
namings resolve: darktable's `IMG_0001.CR3.xmp` names its file exactly,
Lightroom's `IMG_0001.xmp` names the stem, and under the stem the JPEG
beside a RAW is the same photograph and takes the same document, as
DarkRoom's own sidecar already does. Each is reconciled with the catalog
winning — keywords union, a rating or label taken only where the catalog
has none — because a standard XMP carries nothing that could say whether
its value is newer. A genuine disagreement is not resolved; it is written
to a table, and the settings page offers the sidecars' values against it.
That button is the reload the requirement asks to be offered, and the
ETag that moved is the detection it asks for: an `.xmp` edited elsewhere
is exactly a file the pull's ordinary incrementality re-reads.

Writing goes the other way behind a setting that starts off, since NFR-R4
makes writes beside somebody's originals theirs to switch on. With it on,
a judgement or a keyword rewrites the sidecar of whichever spelling
exists, or creates Lightroom's. The record is read from the catalog
whole at that moment rather than carried from the gesture, so a rating
and a keyword a second apart are two writes of one file that agree. And
the file's own title, caption, copyright and hierarchy come through the
rewrite: the catalog has no columns for them, `rewrite` replaces the
owned set wholesale, and a record that said nothing about them would have
deleted them from a Lightroom sidecar on every star.

The rating's two axes cross the format's one field both ways: a
rejection is Adobe's `-1` and stars are stars, and stars arriving on a
rejected frame lift the rejection, since the file said it was worth a
number. An unrated file says nothing and clears nothing, on the rule the
`.drsc` merge keeps. `versions.label` finally has a reader and a writer,
with the code table moved out of the query so the two cannot drift.
2026-09-12 01:08:11 +02:00
dtourolle d3b6127db6 Let a photographer name the state they liked, and go back to it or look at it
FR-DEV-5 asked for named snapshots of an edit state and FR-DEV-7 for a
comparison against a chosen one, and neither existed. The history stack
is per sitting and forgotten with it, on purpose — the gap that mattered
was an automatically saved mis-drag with no way back, and that was closed
first. What was left was the other half: a state the photographer wants
to keep *because* it is worth keeping, which is a different thing from a
step and is not served by making the steps last longer.

A snapshot is an edit state, and an edit state is exactly what a sidecar
version stores, so it is stored as one: a `[version]` block carrying
`snapshot-of = <uuid>`. The parameters, the masks and their parts, the
repairs and the film all arrive through the blocks that already carry
them, a merge keys on the uuid as it does for any version, and a build
that predates the key reads the block as a named version and keeps it —
the right failure. Only the pointer is new. The one reader that has to
know is `default_version`, which must never answer with a snapshot: a
file whose edit is missing is not a file whose edit is one of its saved
moments. The snapshots of an edit are listed by that pointer, oldest
first, the same on every device.

Writing them back removes what this sitting deleted and puts in what it
holds, and leaves standing whatever it never saw — a snapshot the other
device took since the photograph was opened here is not this device's to
remove by not knowing about it. That is the rule the version merge
already keeps, applied one level down, and it is why the save carries
the deleted ids rather than replacing the list wholesale as the masks
are. Each is re-pointed at the uuid the save settled on, because the
default may have been fused onto its canonical identity since the
snapshot was taken.

Restoring is one history step, so undo takes it back whole, as a paste
is. Taking and deleting are not steps: they change nothing about the
photograph, and an undo that removed a snapshot would be undoing a
decision to remember. Holding the eye beside one renders the snapshot
and hands the edit straight back — the same suspension "Before" uses,
against a point the photographer chose rather than the file. Two
sessions on the same photograph get ids that cannot collide, stamped
with the second and a random word, because the merge folds equal ids
into one.
2026-09-12 01:08:10 +02:00
dtourolle 369eb8fbf0 Put the log and the crash records in one file, and show it before writing it
NFR-OPS-1 asks for a diagnostics bundle — the log, the schema version, the
GPU and driver, the app version — "with an explicit preview-and-consent
step before anything leaves the device". The log and the crash records
have existed since August; what did not exist was any way to hand them
over that was not `adb pull` and a knowledge of where the state directory
is, which on the tablet the requirement was written for is nobody.

Nothing here sends anything, and that is the design rather than a gap:
crash.rs already says why a transport built ahead of the consent is the
shape of thing that gets switched on by default. The bundle writes one
text file to a place the user can find, so that they can attach it. That
is the moment it leaves, and it is theirs. So the consent guards the
write, not a send. Preparing gathers everything into memory and shows
what would be written — each section, its size, what was taken out, and
where the file would go — and only the second press puts bytes on disk.
A user who reads the preview and presses the other button has changed
nothing anywhere. The gathered bundle is held between the presses so what
is saved is exactly what was shown, not a second gathering that differs
by whatever was logged while they were reading.

One text file rather than an archive, because a `.txt` opens wherever
the user is sitting and pastes into an issue, and because the preview
can then be the file rather than a summary of it. Every line goes
through the blunter of the two redactions on the way in, whatever the
sink already did to it: the log's own rule keeps paths, since a path
read over `adb` is context, but a file meant to be attached to a public
report by someone who may not read it first is held to the crash
record's rule instead.

The About page's graphics line gains the driver, which the requirement
names and the adapter has always reported. And docs/outstanding.md is
corrected on both OPS requirements: it said crash reporting was a
log::error! hook and NFR-OPS-1 had nothing behind it, and neither had
been true since 2026-08-30.
2026-09-12 01:08:10 +02:00
dtourolle 4574c35236 Let a part be left out of a mask without being taken out of it
A layer built from parts was missing the one control a correction most
often wants: seeing what it did. The question a subtracted gradient
raises is whether it took only the sky, and the question a stroke raises
is whether it filled the shoulder — and the only way to ask either was
to remove the part and look, which answered the question and lost the
part. The layer's own ring answers a different question, about the
adjustment, and hiding eight layers to check one correction is not an
A/B anybody performs.

So a part carries `hidden`. It is an edit and a history step, as the
layer's switch is, and it is folded into the render fingerprint because
hiding a part changes the mask as surely as removing it does. Where the
mask is built the shown parts are walked rather than the parts, which
is what makes a hidden base hand the fold to the first part that is
shown — and a revealed layer whose every part is hidden clears its slice
rather than leaving whatever the last rasterisation put there to be
read back. `covers` asks the same shown parts, so a layer whose only
adding part is hidden costs no slice at all.

In the sidecar the key is `hidden`, in the part's block or, for the
base, in the mask block — under a word that cannot be confused with the
layer's `enabled`, which has always meant the layer. Absent means shown,
so no file written before the switch existed reads any differently.

The row wears the same ring the layer does, one row down, because it is
the same question about a smaller thing.
2026-09-12 01:08:10 +02:00
dtourolle 9cc52fd72b Bind the two develop gestures that were described and not bound
FR-DEV-16's book said resetting a control and hiding a mask layer were
reachable by pointer and by finger, and stopped there. The reason was
honest: the generated rows have no focus, so "reset the focused control"
named a thing the panel could not point at. But a photographer at the
keyboard means something narrower than focus. They mean the slider they
just dragged too far, and that is a thing the panel can remember.

So the Adjustments global keeps the last control moved — two indices,
written where the panel forwards the change and cleared when the next
photograph opens, so a reset cannot reach back into the previous edit
through an index that happens to be shared. R puts it back, through the
same callback the track's double-click takes, and is silent until
something has moved.

The mask layer needs no such notion, because the panel already has a
selection: the rows the edge controls point at. H hides or shows those,
through the path the ring at the head of the row takes, so it is an edit
and a history step exactly as the ring is. A mixed selection goes to
shown, since the layer nobody can see is the one being asked about.

Both tags now carry the key, and the book says so.
2026-09-12 01:08:09 +02:00
dtourolle 2836ec2881 Build the Windows installer in a container, and run it under Wine
docs/windows.md specified it; this is §9 steps 1, 2 and 4 run, and the
report in §10. A Debian trixie image with rustup, the MinGW cross
compiler, NSIS and Wine; a build.sh in the shape of the Android one;
a package.sh that stages the executable and the seven models behind
the same LFS-pointer guard every other packager carries, then runs
makensis; and the .nsi itself — per-user, no elevation, an uninstaller
that leaves the library alone.

Measured: the executable links first time once the link flags were
right, imports only Windows system DLLs, prints its version under
Wine, and the installer installs and uninstalls silently under Wine
with the registry key and the models where §5.2 says. What Wine
cannot show is the Start Menu shortcut: CreateShortcut is IShellLink
and does nothing headless.

Four claims in the spec's first draft were wrong and are corrected in
place with the reasoning kept: the whole-archive winpthread flag
breaks the link and was never needed; build scripts need a host gcc;
bookworm's Wine lacks the bcryptprimitives.dll rustc's std imports,
so the image is trixie; and NSIS's default stub is 32-bit, so the
installer says amd64-unicode and needs no i386 Wine.
2026-09-12 00:54:11 +02:00
dtourolle fa4dca327f Give the desktop executable a version flag and a Windows identity
Three things the Windows build showed the entry point was missing, and
that a Linux build never asks for.

`--version`, answered before the logger and the crash hook install: a
binary built on a machine that cannot run the application — the Linux
CI producing the Windows executable, checked under Wine — needs an
exit that proves it starts without opening a window or touching the
user's directories. It is the smoke test in docs/windows.md §6.

A GUI-subsystem executable in release, or Windows keeps a console
window open behind the application for the life of the process. Debug
builds keep the console, which is where their log goes.

A resource block, or Explorer, the Start Menu and the taskbar show the
generic executable icon and the Details tab is empty. build.rs wraps
the PNG every other platform uses into an .ico at build time — an ICO
entry may be a PNG, so the wrapper is a 22-byte header — and hands it
to winresource with the version cargo already knows. The crate is an
unconditional build-dependency because a cfg(windows) on one is
evaluated against the host, which here is Linux; the script itself
returns before touching it on any other target.
2026-09-12 00:54:11 +02:00
dtourolle 0a2c49dd10 Guard two constants the Windows target leaves unused
secrets.rs names the keyring service and desktop_client.rs the socket
timeout, and every use of both sits under a cfg that a Windows build
does not satisfy — the placeholder secret store has nothing to file
under, and the Nextcloud client's named pipe is not opened yet. The
first cross-compile reported both as dead code, which is a failed
clippy job the moment the Windows leg runs with -D warnings. Guarded
by the same cfgs as their users, with the reason beside each.
2026-09-12 00:54:10 +02:00
dtourolle b718c70b11 Specify a Windows installer built by the Linux CI
The tree is closer to Windows than a Linux-only project usually is:
every image library, the TLS stack and the inference engine are pure
Rust, and dr-plat already keeps the Linux-only code behind cfgs with a
loud fallback where none exists for another platform. What remains is
a short list above dr-plat — five XDG path lookups, an xdg-open, the
secret store's third implementation, the models' lookup beside the
executable — and none of it touches core, which is the NFR-PORT-3 test
this would be the first real run of.

docs/windows.md decides the GNU target over MSVC-via-xwin, Vulkan only
as on every other platform, a per-user NSIS installer that leaves the
library alone on uninstall, and a CI leg in the shape of the Android
one. It is explicit about what a runner with no Windows can verify —
that it links, is PE32+, starts under Wine and installs under Wine —
and what it cannot, which is everything involving a real GPU driver.
Three FR-PLAT-WIN requirements and a channel row record the decisions;
the ordering puts a first cross-compile on the developer machine before
any container exists, because the list of cfg gaps is a reading of the
source and the compiler's list will be longer.
2026-09-11 23:30:16 +02:00
dtourolle a2c7789007 Sign the Android build with a real key, and let package.sh use it too
Benchmarks / CPU and I/O (per commit) (push) Successful in 2m52s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Successful in 44m13s
Build and test / Layer separation (push) Successful in 56s
🐳 Android image / Build and push (push) Successful in 17m16s
Build and test / android-image (push) Successful in 17m17s
Traceability / Requirement traces (push) Successful in 1m6s
Build and test / Android (aarch64) (push) Successful in 23m37s
The release keystore now exists and its four secrets are loaded into
Gitea, so CI produces an APK a device can update in place. Until now
every build, CI and local alike, was signed with a throwaway debug key
-- CI's fresh per run, the local one exactly as durable as the cache
directory it lived in -- and the night that cache was cleared, no build
anywhere could install over the tablet's copy.

package.sh forwards KEYSTORE_PASS, KEY_PASS and KEY_ALIAS into the
container and copies the keystore under the mounted target directory
for the build, so a local release-signed build is one environment line.
The doc records where the local copy of the key lives.
2026-09-11 23:22:28 +02:00
dtourolle 7c44740d9f Skip the read-only-directory test where modes are not enforced
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m41s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Successful in 1h44m9s
Build and test / Layer separation (push) Successful in 48s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Traceability / Requirement traces (push) Successful in 36s
Build and test / Android (aarch64) (push) Successful in 1h1m51s
`a_failed_overwrite_puts_the_original_back` makes the presets directory
read-only and expects the overwrite to fail. CI's Desktop job runs in a
container as root, and root is not refused by a mode: the write succeeds,
the assertion fails, and build-and-test has been red on every push to
master since the test arrived.

The test now probes the refusal it depends on -- one write into the
directory it just locked -- and skips where that write goes through.
Probed rather than keyed on the uid, because what the test needs is the
refusal itself, and a filesystem mounted without permission checks would
pass a uid test and fail this one all the same.
2026-09-11 22:22:46 +02:00
dtourolle 4f31123b0c Let the user choose which SCRFD finds their faces
faces.md §12.3 measured what the cheapest detector costs: the small
faces in every group shot, and a dog embedded a dozen times. Which
trade is right depends on the machine doing the sweep — a desktop left
overnight and a tablet on a battery want different answers — so the
detector is now a per-device setting, Fast / Balanced / Thorough on
the settings page beside the indexing button, persisted with the rest
of the settings file.

A detector is half of a model id. Every face, marker, shard and
calibration is keyed on faces.model_id precisely so that a model change
is a new id and a re-index rather than a silent change under existing
data, and a detector change is a model change: it decides which faces
exist and where the landmarks that align them land. So each choice
names its own pipeline. 500M keeps the bare "w600k_mbf" every existing
library was written under, so an upgrade disturbs nothing; the others
are qualified. Choosing one restarts coverage from zero under the new
id, the sweep re-detects, confirmed names carry across by box overlap,
and the sync shards are keyed by the same id so a peer on another
setting neither adopts nor pollutes them. The library controller
carries the id into the sync the same way it carries the cache budget,
because the sync starts from places that have no settings in reach.

All three shape-fixed exports ship — APK, Arch, Flatpak — since a
tablet has no other way to obtain the one it was not installed with;
the APK grows by twenty megabytes for the choice.
2026-09-11 22:12:53 +02:00
dtourolle adf5d6cdd9 Drop a rival pipeline's marker when an image is re-indexed
record_detections replaces every face on an image whatever model found
them, but left the other models' face_index rows standing. With one
model that was unobservable. With a second pipeline it leaves an image
marked "done" under the first with none of its faces behind the marker
— the state the V12 repair existed to undo — and a user who switched
back would find those photographs permanently empty.

An image now holds the faces of whichever pipeline looked at it last,
and only that pipeline's marker. Confirmed names still carry across by
box overlap, since they were read before the replacement.
2026-09-11 22:12:40 +02:00
dtourolle 9d35addd86 Measure what the cheapest SCRFD actually costs in faces
§1 chose scrfd_500m on FLOPs and never measured the recall it gave up.
A dr-ui example now runs several detectors over the same sample of
stored proxies, matches boxes by IoU against the first, buckets the
result by face size, times each, and writes contact sheets of the
disagreements in both directions — because a count of extra faces says
nothing until someone has looked at whether they are faces.

Over 400 proxies from the reference library: 2.5G finds 14% more faces
for 12% more time, 10G a further 12% for 3.1× the time. The extras are
small real faces. The 86 faces only 500M found are a dog a dozen times,
a stop sign, a wheel and the backs of heads. Recorded in faces.md §12.3.
2026-09-11 22:12:39 +02:00
dtourolle 3d6d69ec90 Wrap the face-sweep repair match the way rustfmt wants it
Benchmarks / CPU and I/O (per commit) (push) Successful in 4m5s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 39m17s
Build and test / Layer separation (push) Successful in 1m24s
🐳 Android image / Build and push (push) Successful in 10s
Build and test / android-image (push) Successful in 10s
Traceability / Requirement traces (push) Successful in 55s
Build and test / Android (aarch64) (push) Successful in 27m8s
CI's Desktop job failed at the Format step on 16f3fb4: rustfmt puts the
`match faces_without_proxy(...)` on its own line under the `let` and
re-indents its arms, and the commit was written without running it. No
code changes; only the layout of that one match in library.rs.
2026-09-11 22:00:38 +02:00
dtourolle 16f3fb41a3 Measure the faces already found rather than finding them again
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m13s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 54s
Build and test / Layer separation (push) Failing after 1s
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 2s
Traceability / Requirement traces (push) Successful in 52s
Build and test / Android (aarch64) (push) Successful in 30m6s
Every face stored before its quality was kept holds a unit vector, and
V14 forgot the run marker of each image holding one so that the next
sweep would look again. Looking again meant detecting again: a whole
re-detection per image, with every suggestion on it thrown away and the
confirmations carried across by box overlap, to recover one number.

The sweep now has a measuring pass between the proxy repair and the
un-indexed images. It lists every image holding an unmeasured face,
fetches the original once, warps each stored face from the landmarks it
already has, embeds it, and writes the raw vector and its length over
the old row. Ids, boxes and identities are untouched; the marker is
re-written fresh so the sync exports the measured vectors. A face whose
landmarks no longer make a warp is dropped, as detection would have
refused to store it. `faces_unindexed` leaves those images to the
measuring pass, so the V14 deletion no longer costs a second detection.
2026-09-11 21:50:12 +02:00
dtourolle 8b3abdb787 Keep each face's quality, and never compare against a poor one
The embedder's raw output has a length, and the length is a reading of
how recognisable the crop was: a blur, an occlusion or a hard profile
comes out short. Normalising threw it away. A short vector sits near
the middle of the sphere and matches a little of everyone, which is how
one bad crop bridges two people in a grouping pass.

So the length is kept — the store now holds the raw vector, re-normalised
on load, with the length beside it as `faces.quality` — and a face under
MIN_GALLERY_QUALITY (14) is a probe: measured against the gallery and
placed where it fits, but never what another face is measured against.
Two probes are never paired, and a probe is nobody's evidence for a
confidence. The People screen shows the number as "Quality 17.3", dimmed
below the floor.

Faces indexed before this stored unit vectors and have no reading; they
are admitted to the gallery, and schema V14 forgets the run marker of
every image holding one so the next indexing pass measures them. A
peer's unmeasured shard faces are not adopted, or a sync would write
that marker back.
2026-09-11 21:50:12 +02:00
dtourolle a87139b838 Give every mask an eye and a colour, and put the brush where the mask is
The first build of seeing a mask showed the selected layer's, in one global
style, from a strip at the top of the panel. It answered the wrong question and
answered it somewhere nobody looked. What a photographer asks of two masks is
how they meet — where the sky's edge sits against the building's — and that
needs both on screen at once, in colours that can be told apart.

So each row of the stack has an eye, drawn in the colour its mask is shown in,
and each mask has six swatches to choose that colour from. Several can be open
at once; a new one comes up open, in the first colour nothing else is using.
The style — tint, alpha, outline — is the one setting that stays global, above
the stack, because three styles at once are three pictures that cannot be read
against each other. Alpha now draws every shown mask, each in its colour, on
black. In the pipeline a `Reveal` is a list of `(layer, colour)` rather than
one layer, and every reveal block carries its own colour.

The brush moves too. Select, Paint and Erase and the three sliders under them
sat at the top of the panel, appeared only once a row was selected, and said
nothing about which mask they acted on — so "how do I paint" and "how do I
correct the model's outline" both had the same answer and nobody found it.
They sit under the selected mask's parts now, beside the swatches, and on a
subject or a category the hint says what a stroke there does: it becomes a
part of this mask, joined to the model's, and can be taken out again.

Eyes and colours are viewing state, on the session and not on the layer, so a
photograph reopened has every eye closed — the stored-mask round-trip test
asserts it.
2026-09-11 19:03:51 +02:00
dtourolle 936490880b Release 0.12.0
Benchmarks / CPU and I/O (per commit) (push) Successful in 12m48s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 37m37s
Build and test / Layer separation (push) Successful in 46s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 3s
Traceability / Requirement traces (push) Successful in 1m4s
Build and test / Android (aarch64) (push) Successful in 53m59s
2026-09-11 09:33:36 +02:00
dtourolle d8b9b5a4bb Let the release script write the release commit it was never trusted with
Every release commit in the history reads `Release X.Y.Z` and none of them
carries the message this script would have written, so nobody has ever
passed it `--commit` — and the reason is in the message it wrote: a
Co-Authored-By trailer naming an assistant, which no commit in this repository
carries and none should.

The trailer goes, and so does the paragraph above it: the script's own header
already says why it exists, and a release commit is the one place a one-line
subject is the whole convention.
2026-09-11 09:33:28 +02:00
dtourolle ac0aea70ec Show a mask as soon as it is made
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m52s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 39m53s
Build and test / Layer separation (push) Successful in 1m0s
🐳 Android image / Build and push (push) Successful in 4s
Build and test / android-image (push) Successful in 4s
Traceability / Requirement traces (push) Successful in 46s
Build and test / Android (aarch64) (push) Successful in 25m19s
Choosing a category is asking what it selected, and for a subject or a
category that question had no other answer on screen: the model's outline is
not derivable from anything visible, a fresh layer carries no adjustment to
judge it by, and the list it was chosen from says "architecture 23%" without
saying which 23%. The control that draws the mask existed but had to be found
and pressed, a panel's height away from the list the choice was made in.

So a new layer arrives with its mask showing, from the resting position only.
Somebody who has chosen the alpha or the outline keeps it, and nothing re-arms
in the background — every caller is a press that asked for a new mask.

That makes the canvas depend on how a layer arrived, which is correct and
worth stating: a session that has just made a mask draws a frame that a
session which read the same mask out of a sidecar does not. Viewing state is
not edit state and does not travel in a file, and
`a_stored_mask_renders_exactly_what_the_model_rendered` now says so at both
ends.
2026-09-10 20:56:11 +02:00
dtourolle 5d175cc668 Let a press on the photograph reach the tool that was armed for it
The brush did nothing, and neither did three other things nobody had tried
lately: clicking a subject on the photograph to select it, placing a repair,
and sampling a neutral. All four are TouchAreas over the canvas, and all four
sat behind the pan/zoom area, which is full-canvas and enabled for everything
but a crop. It took every press in the viewport and they were never offered
one.

Slint hit-tests siblings front-to-back (`send_mouse_event_to_item` visits
children `TraversalOrder::FrontToBack`), a TouchArea answers `GrabMouse` on any
press it is enabled for, and the first grab aborts the traversal. Front means
*last declared*. Each of the four carried a comment saying it sat "above the
pan/zoom area so a click reaches it first" — true of the order they were
written in, and backwards.

Nothing about the geometry decides this, so nothing about the geometry could
have fixed it. The pan area is declared first now, as the backstop it always
meant to be, and the rule it leaves behind is that the general case goes above
the specific ones. `GradientHandles` is the other end of that rule and is why
dragging a handle has worked all along while everything between it and the pan
area did not.

The order is asserted in a test, because this is a fault that compiles, passes
every other test, and silently removes four tools at once.
2026-09-10 20:56:10 +02:00
dtourolle 76ad667fd6 Offer a mask that is nothing but a hand
Every route to a layer began with a selection — a gradient, a band, a subject,
a category — and painting was reachable only by making one of those and
joining a painted part to it. So the answer to "brush a correction onto this
corner of the sky" was "add a radial gradient you do not want, then paint into
that", which is not an answer.

Paint sits beside Linear and Radial and makes a layer whose base is a brush.
It covers nothing until a stroke lands in it, so pressing it arms the brush
and shows the mask as well: a row that appeared and changed no pixel, with the
pointer still in "select", is indistinguishable from a button that did nothing.
2026-09-10 20:28:09 +02:00
dtourolle c045702a47 Show the photographer the mask they are shaping
Nobody can refine an edge they are not being shown. The only thing drawn on
the canvas was the region overlay — a false-coloured picture of what the model
*detected* — which knows nothing of a layer's feather, its falloff, its
morphology, its invert or its opacity, and nothing at all about a gradient, a
range or a stroke. Every control added for mask editing therefore acted on
something invisible, which is why the whole feature reads as absent rather
than as unfinished.

A layer's finished mask now draws over the photograph in one of three styles:
a tint for whether the right thing is selected, an alpha for where the edge
is, an outline for whether that edge is registered against the detail the
other two hide.

The hard part is not the shader. A selection with no adjustment on it changes
no pixel, so it is not active, so it holds no slice of the mask array and is
never rasterised — and that is exactly the layer somebody wants to look at,
for the whole of the time between choosing a subject and deciding what to do
to it. So `MaskStack::rendered` is `active()` plus the layer being looked at,
and the rasteriser, the composer and the distance-field builder all index by
position in it. Which is also why the design's "two uniforms, no recompile" is
not available: a uniform can select a slot, it cannot conjure one.

The reveal is never on the graph. It reaches the pipeline as an argument to
`compose_revealing`, and `compose_for` — which the exporter, the thumbnail and
the neutral probe all call — has no way to ask for one. A flag on the graph
would have been shorter, would have type-checked, and would have been one
forgotten reset away from a red tint baked into an exported file.

And the tools that shape a mask now arm. `Masking.tool` is an `in` property
only Rust may write, and the handler wrote nothing back, so the strip reported
"Select" however many times Paint was pressed and the paint area was never
enabled — the brush, the parts and the whole of FR-DEV-19b reachable from no
control in the application.

The region overlay stands down while a mask is being shown, and its button now
says what it hides: two overlays that look alike and mean different things is
worse than either.
2026-09-10 20:28:06 +02:00
dtourolle 193b35a249 Start a category mask where the photograph can bear it
Clicking "architecture" made a layer whose mask was gone. Every category layer
began at STRICTNESS_DEFAULT, and that constant was fitted on the synthetic sky
the refine tests build — its own note warns that a real photograph's noise
"moves every crossing down together", which turns out to be a considerable
understatement. Measured over seven ordinary frames, half scale removes 76% to
99.5% of `architecture`, 36% to 93% of `ground` and 18% to 91% of
`vegetation`. Only sky, the category the number was calibrated against,
survives it.

An empty mask is indistinguishable from a broken one: the layer is listed, the
adjustment moves, and no pixel changes. So what this looks like from outside
is that the segmentation does not make masks at all.

No smaller constant fixes it either, because a nat of evidence means different
things over a smooth sky and over a stone facade — the useful position is
above 5 on one frame and below 1 on the next. So the frame is asked instead:
`Refinement::gentle` walks down from half scale and takes the first rung whose
gate removes no more than a sixth of the category's weight, and the model's
own outline when none of them does. One `apply` on a friendly photograph and
four on an unfriendly one, paid when a layer is made rather than for eight
categories nobody masked.

The slider's reset went to 4 as well, so taking the control back to its
"default" emptied the mask. It goes to zero now, which is the one position
documented to mean something: exactly what the model weighted.
2026-09-10 20:27:40 +02:00
dtourolle 404fea47a8 Wrap the lines the merge resolution left long
Benchmarks / CPU and I/O (per commit) (push) Successful in 2m59s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 33m34s
Build and test / Layer separation (push) Successful in 55s
Traceability / Requirement traces (push) Successful in 42s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Successful in 23m43s
`cargo fmt --check` failed the desktop job, on three files and for one reason:
routing the mask handlers through the `Masking` global was done by substituting
the call prefix, which is a text edit rather than a Rust one. It left
`window.global::<Masking>().on_part_join_picked(...)` on a line that had been
short enough as `window.on_mask_part_join_picked(...)` and no longer was.

Formatting only. The whitespace-stripped source is identical in the two `ui/`
files; the third differs by the trailing commas rustfmt adds when it breaks a
call across lines.

The matrix moves with it, because the tags shift by a few lines and the check
compares line numbers.
2026-09-08 08:57:32 +02:00
dtourolle d920716a2b Release 0.11.0
Benchmarks / CPU and I/O (per commit) (push) Successful in 11m47s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 37s
Build and test / Layer separation (push) Successful in 38s
Traceability / Requirement traces (push) Successful in 1m0s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Build and test / Android (aarch64) (push) Successful in 55m56s
2026-09-07 22:31:53 +02:00
dtourolle 2d878c2117 Offer the film stock in its own group, and let its list scroll itself
Two faults in one control, both reported from the tablet.

The stock picker appeared in every group. It is not a parameter, so it is not
a row, so the filter that hides every other control when a group is chosen
never saw it — "Kodachrome" sat at the top of Light, of Colour and of Detail
alike. Three places it does not belong, and the one it does no more prominent
than the rest. The descriptor has said `Effect` and only `Effect` since the
film moved there; nothing was asking it.

So the panel now asks. It cannot ask directly — a generated panel may not know
which operation a control belongs to — so the session answers, from what the
operation declares it is about, and a stock re-declared as something else would
move on its own. The flag is recomputed when the group changes as well as when
the film does, which is the half that would have made it stale exactly when it
mattered.

And the open list was unbounded, so it made the develop column taller and the
column scrolled as one: reaching Velvia dragged every slider below it off the
screen, an answer given once pushing aside the controls used constantly. It now
scrolls within a bounded height of its own.

That viewport is counted rather than measured, for the reason the tool rail
records a few files away: a viewport that asks a layout how tall it wants to be,
while the layout takes its height from the viewport, is a cycle Slint settles by
handing back the height it was given — and the content is then clipped in
silence rather than scrolling. Every row here is one fixed height, so
multiplying is exact.

The group rule has a test. The scrolling does not, and cannot: it is a layout,
and a layout fault is invisible to the compiler and to every assertion that can
be written about it.
2026-09-07 20:42:15 +02:00
dtourolle 4c217c9be6 Show what a control does to a photograph, one parameter at a time
Benchmarks / CPU and I/O (per commit) (push) Successful in 2m52s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 56s
Build and test / Layer separation (push) Successful in 37s
🐳 Android image / Build and push (push) Successful in 4s
Build and test / android-image (push) Successful in 4s
Traceability / Requirement traces (push) Failing after 54s
Build and test / Android (aarch64) (push) Successful in 26m29s
A node that has just been declared can be read, reasoned about and tested, and
none of that answers the question a photographer asks first: what does moving
this do to the picture. Colour grading, dehaze and the range masks were all
argued into the tree on their behaviour and none of them had been *looked* at.

So two diagnostics. `sweep` walks one operation from its minimum to its maximum
and writes a frame per step; `rangesweep` does the same for a mask band, which
is not an ordinary parameter — it lives on a layer, is rasterised by its own
pass, and only becomes visible through whatever adjustment the layer carries,
so it gets two stops down to make the selection legible.

`sweep` names no operation. The id arrives as a string and the parameters and
their ranges come from the graph's own capabilities, so a node declared
yesterday sweeps on the same terms as one that shipped a year ago — the
property `ops/README.md` promises, used rather than asserted.

Three things it learned the hard way and now records. It renders through
`render_detailed` unconditionally, because the fused path refuses a shader
composed with a detail stage rather than rendering it wrongly, and that call
falls through when there is no such stage. It takes `SWEEP_HOLD`, because a
parameter grouped under one widget is not meaningful alone: a hue with no
strength behind it renders the same frame every time, which reads as a broken
node rather than a correctly declared neutral. And it bounds the output, since
a 25 MP frame is a 75 MB PPM and a sweep is hundreds of them.

Both read a rendered file as readily as a raw one, so a JPEG can stand in where
no raw is to hand — on the terms `from_rgba8` documents, with the controls
still working and their neutral being what the camera left rather than what the
sensor recorded.
2026-09-07 20:01:08 +02:00
dtourolle e44cb8cbe0 Regenerate the traceability matrix for the rebased line numbers 2026-09-07 20:01:01 +02:00
dtourolle 0a5eab0487 Wire the develop panels through globals, so a second copy is one line
Every panel in the develop column declared its inputs and its callbacks and
had `app.slint` bind each one to a property or a callback on the window root.
That is fine while a panel is drawn once. N9 draws them a second time, in the
portrait dock, and the wiring is what would have to be copied: `MaskPanel`
alone ran to forty lines of forwarding, and a callback added to one copy and
not the other compiles, renders, and simply does nothing on the layout nobody
was looking at.

So the wiring moved to Slint globals. A panel reads the global and calls the
global; Rust hooks the global instead of the window; and the instantiation in
the column is now the panel's name and a pair of braces — every one of the ten
children of the column, with no property that differs by placement left to
supply.

There is a global per panel family rather than one for all of them, and the
reason is an import cycle. Each panel's model struct — `ParamRow`, `MaskRow`,
`HistogramView` — is declared in the panel's own file, so a single global
holding `[MaskRow]` and `[ParamRow]` would have to live in a file importing
`masks.slint` and `adjust.slint` while both imported the global back, which
Slint rejects. Breaking that needs six model declarations relocated, which is a
change to the data model and not to the plumbing this is about. A global beside
the panel it serves also lets each name drop the prefix it was carrying only
because the window root is one flat namespace: `root.spot-radius` is
`Repair.radius`, and `root.peaking-on` is `Peaking.showing`.

`session.slint` is new and holds the two facts every family needs and none of
them owns: whether there is an open photograph to edit, and which mode the view
is in, with the three readings of the mode derived once instead of at each of
the dozen places that tested one. `ViewMode` moves there from `adjust.slint`,
where it was only ever a lodger.

Nothing on screen changes. What is not here: the tool rail and the status strip
still take their properties at the instantiation, because they are drawn once
and N9 does not copy them; the preset sheet's own state stays on the window,
because the library grid opens the same sheet and a global cannot bind the
window's state — which is why `Transfer.open-presets` is handled in
`presets.rs`, beside the summary it already had to compute.
2026-09-07 20:01:01 +02:00
dtourolle dd14243dba Lay the develop view out by coordinate, so the column can sit below
Slint cannot turn a layout on its side, and that is what D-N7 asks for. So
the HorizontalLayout holding the rail, the canvas and the develop column
becomes a plain Rectangle and each of the three states its own x, y, width
and height. With `column-below` false those come out where the layout put
them to the pixel — a HorizontalLayout has no spacing or padding of its own,
the rail and the column took their declared widths at the two edges, and the
canvas was the only child that stretched.

The alternative was the column subtree declared twice under two `if`s, which
is four hundred lines of bindings copied, in a file whose own notes record a
conditional child in a layout as the shape that has produced binding loops
here before.

The column is the same column either way: the same contents, the same
Flickable, the same toggle. `panel-visible` collapses the dock's height
exactly as it collapsed the column's width, so `column-width` and
`dock-height` are each zero unless the column is both open and on that axis,
and the canvas can subtract both without asking which case it is in. The
stack inside stretches to the dock's width on its own — a layout that is the
direct child of a Rectangle fills it, and the Flickable's viewport was
already bound to its own width. Nothing is reflowed; N9 does that.

`dock-height` is mandated in style.yaml for the reason `panel-width` beside
it is, on the other axis: a dock that sizes itself to its contents is a
photograph that changes height when a caption wraps. 480 until N6 measures
the device.

The seam follows the column round: a hairline down its left edge beside the
photograph, along its top edge under it, so it stays between the two.
2026-09-07 20:00:50 +02:00
dtourolle 5f0b11c1f4 Ask the window how tall it is, and say when the column belongs below
D-N7 puts the develop column under the photograph on a tall window, and the
axis it turns on is aspect rather than width: a 960-wide portrait tablet is
expanded by width and wants the dock, a 1500-wide landscape desktop is
expanded by width and does not. So this cannot be folded into the layout
class, and it is not remembered per class either — closing the column in
landscape closes the dock in portrait, because it is the same column.

`window-resized` reported width alone and now reports both, from a
`shell-height` that subtracts the safe-area insets exactly as `shell-width`
subtracts them: on Android the strips the status and navigation bars occupy
are on the axis being measured, so the aspect of the window and the aspect of
the space the interface actually gets are not the same number.

`column_below` is the decision, with two thresholds rather than one. It is
read on every resize event, and a single threshold means a window dragged
along its own diagonal crosses it several times a second while the pointer is
still down. Entering at 1.25 and leaving at 1.15 is a dead band no plausible
drag re-crosses.

The comment on EXPANDED_MIN_WIDTH claimed a tablet in portrait gets the
compact layout. It does not — its panel is about 960 logical pixels across,
which clears 820 — and that mistaken example is the one D-N2 reasoned from.
Corrected in the same breath, since this is the commit that says what
portrait actually changes.
2026-09-07 20:00:50 +02:00
dtourolle 30468c4c69 Put the window-metrics doc on the function it describes
The comment explaining why both coordinate systems go on one line was
written for `log_window_metrics` and sat above `window_metrics_level`,
where it read as the start of that function's much longer note. Two
doc blocks ran into each other and the one that prints had none.
2026-09-07 20:00:50 +02:00
dtourolle 428d8c4a51 Say what size the window actually is, in both coordinate systems
Every figure in D-N7's table was computed at a guessed scale factor. The
tablet's panel is 3000 by 1920 physical and nothing in this repository
has ever recorded the density Android reports for it, so the dock's
width is either 900 or 1037 and its available height is 200px either
way. N6 asks for the measurement; this is the line that carries it.

Beside the existing `apply_layout_class` call, because that is where the
window is already being asked for its size and its scale, and again on
every resize, so turning the tablet over records the other orientation
in the same logcat. Both coordinate systems on one line: a logical size
cannot be checked when the scale is the thing in doubt, and a physical
size that does not divide by the scale printed next to it says the
reading is of something other than the panel.

The level is not fixed, because none of the three obvious choices works.
`android_main` caps the facade at info, so debug never leaves the
device and a debug-only line answers nothing. A drag emits a resize per
frame and each accepted record is also appended to the on-disk log, so
info on every resize is not a diagnostic. And the first reading is not
the settled one: on X11 the window reports 0x0, then 360x320 at scale
1.0, then 1100x720 at scale 2.0, so reporting only the first would put a
number in logcat that is not the window's.

So the pair that decides is the scale factor and the orientation --
exactly what N6 is asking for, and exactly what a drag leaves alone. A
window with no area is not a reading and records nothing. Every later
change to either half, the scale resolving or the tablet turning over,
is a new answer and goes out at info; everything else is debug.
2026-09-07 20:00:49 +02:00
dtourolle 2361b2d4ef Tablet portrait is the expanded layout class, not the compact one
FR-UI-1's table has put tablet portrait in the compact class since the
register was written. D-N2 settled that it does not belong there: a
12-inch tablet is about 1024 logical pixels across in portrait and
EXPANDED_MIN_WIDTH is 820, so both orientations of both targets are
expanded, and compact fires only on a desktop window dragged narrow. The
code has always agreed - apply_layout_class reads the window's width and
nothing else - so the register was the last place still saying otherwise.

What portrait actually needed was a different axis. D-N7 docks the
develop column under the photograph on a tall window, set from the
window's aspect and independent of the layout class: a 960-wide portrait
window is expanded and wants the dock, a 1500-wide landscape one is
expanded and does not. That is a placement rather than a mode, so the
clause about the transition being continuous stands as written, and the
requirement now records the side the column takes as its own property.

FR-UI-2 gains one clause for the same reason. It said modality affects
control sizing and affordances "not layout", which D-N6 reversed: under
touch the group selector leaves the strip above the column for the tool
rail, and groups-in-rail in toolrail.slint is what moves it.
2026-09-07 20:00:40 +02:00
dtourolle de32c04a68 Dock the develop column under the photograph on a tall window
D-N2 dismissed portrait with one number: a 12-inch tablet is about 1024
logical pixels across, which clears the expanded breakpoint. That was
worked out for a 4:3 panel. The tablet's is 3000 by 1920, and on that
aspect a column beside the photograph in portrait leaves it a strip 540
wide and 1456 tall: a 3:2 frame gets 540 by 360 where a column below it
would give 900 by 600, and the portrait frame gains too.

D-N7 records the decision: a third property beside the layout class,
derived from the window's aspect with hysteresis, that lays the same
rail, canvas and column out on the other axis. Not a sheet, not a second
layout, and the rail does not move. N6 measures the device the numbers
were guessed for, N7 does the frame with the stack stretched as a
stopgap, N8 moves the develop callbacks onto a global so the panels can
be declared twice cheaply, and N9 is the three-column composition the
dock's width is actually for. N5's "no strip along the bottom" is struck
where D-N7 reverses it and kept where it does not.
2026-09-07 20:00:40 +02:00
dtourolle e13d3a54fc Watch the mask tools work, rather than reading that they do
A still frame cannot show what makes these tools right or wrong. What
matters is how the mask *moves*: whether a stroke lands where the finger
went, whether a subtraction takes away only what it covers, whether an erase
inside a correction punches through the selection underneath. Every one of
those is a sequence, and the test suite asserts single pixels.

So this renders the sequences. A synthetic photograph, one frame per step of
each mode — painting, erasing, joining a part and taking it out again,
inverting, and sweeping the edge controls — as PPM, which ffmpeg turns into
a GIF in one line. It runs headless, needs no RAW and no model, and takes a
few seconds.

It is also the honest answer to "show me it working" while the tools are
still being wired to a finger: this is the pipeline itself, not a mock-up of
it, and a fault in the fold shows here as a frame that looks wrong.
2026-09-07 20:00:40 +02:00
dtourolle df741a8a49 Let one mask be built from more than one selection, and paint into it
A mask the model draws arrives approximately right — stopping inside a
shoulder, leaking into the hair — and FR-DEV-3's edge controls move the
*whole* boundary, so no value of feather or dilation fixes two errors that
go opposite ways. What fixes them is a second selection joined to the first,
and a layer that held exactly one source had nowhere to put one. The brush
the core has had all along was reachable from no control in the application.

A layer is now an ordered list of parts. Each names a source and how it
joins the mask before it — added to it, or taken out of it — and carries its
own edge treatment, because a model's soft coverage and a stroke painted
where it stopped short do not want the same feather. Invert and opacity stay
on the layer, where the composed shader already reads them.

The sidecar grows `[part]` blocks and nothing else. A layer of one part
writes exactly the bytes it always did; a mask block with no part blocks
after it reads back as one part; and a stroke, a join or a source this build
cannot read costs that part rather than the layer. So every sidecar in every
library still parses to the edit it always was.

On the device the parts fold into the layer's one slice, so eight layers
still cost eight channels: union is a `max` blend and subtraction is the
erase blend the brush already used. A part is drawn into a scratch texture
before it is joined, and that is not incidental — an erase stroke means a
hole in *that part*, not a hole in the mask, and drawn straight onto the
accumulator it would punch through the subject underneath. A layer of one
part skips all of it and takes the path it always took.

In the interface: a part list under the selected layer with a chip saying
which way each joins, Add and Subtract beside it, a Select/Paint/Erase strip
with the brush's size, hardness and flow, and a drag on the photograph that
paints. Pressing Paint on a mask that cannot hold a stroke joins a part that
can, rather than explaining that a subject is not a brush. A whole stroke is
one step in the history.

The edge controls now shape the part that is selected rather than the layer,
which is the one behaviour change to an existing control: with a correction
selected, the feather slider softens the correction and leaves the model's
mask alone.
2026-09-07 20:00:40 +02:00
dtourolle 9ede23073d Specify the tools that edit a mask once the model has drawn it
The auto masks arrive in a second and cannot then be changed by a pixel: no
brush, no way to cut one selection out of another, no way to drag a boundary
that stopped inside a shoulder, and no way to see the alpha a layer actually
produces — the canvas overlay draws what the model detected, not the mask.

docs/mask-editing.md is how that closes. A layer stops holding one source and
holds an ordered list of parts, each naming how it joins the mask before it,
so painting on an auto mask, subtracting, and intersecting a subject with a
luminance band are all one mechanism. Notes what the tree already has (the
whole brush is written and reachable from no control), what the set operations
actually cost (three blend states, no new texture), why the edge push is a
warp rather than a local morphology, and why painting has to draw
incrementally. Ends with the four decisions the build needs first.
2026-09-07 19:59:44 +02:00
dtourolle 1e171c6d31 Let a collection be picked up, rearranged, and emptied after the fact
Collections could be made and filled and never reorganised. Nesting had
a drag; un-nesting had nothing, in either direction — "All photographs"
refused every drop, which is right for a photograph and wrong for a
collection, which has a top level to be returned to. So a collection put
inside another was in there permanently. Right-click deleted an *empty*
collection outright and refused otherwise, which is wrong in both
directions at once: destructive with no confirmation, and no way at all
to delete a collection that held anything without emptying it by hand,
child by child. And a photograph could only leave the collection the
grid was scoped to, since that is the only one a button in the header
can name — the cell's badge says a photograph is in three collections
and never which three.

Three ways in, one vocabulary:

**Hold a row.** The tree is inside a Flickable, which claims any drag
beginning inside it, so with a finger a drag on a row is a scroll until
something says otherwise. The hold is that something. It lifts the row —
drawn before anything moves, so the gesture says it has been understood
— and then what the user does decides which of two things they meant:
move, and it is a rearrangement; let go, and it is the row menu. The
same fork the grid already uses to tell hold-to-select from drag-to-file.
`decide_release` is that fork, and it is tested, because getting it
wrong one way puts a sheet over every tidied tree and the other way
makes the menu unreachable by touch.

**The row menu.** Rename, new collection inside, move to top level,
keep offline, delete. Deleting asks once when there is anything to lose
and says what survives: the photographs stay in the library, and nested
collections move up rather than going with it — which is what the
catalog does, and what a user would never assume. An empty collection
goes on the first press, because a dialogue about losing nothing is how
people learn to dismiss dialogues.

**"Collections…" on a selection.** Every collection the selection is
filed in, each with a count — "3 of 40", so nobody takes forty
photographs out of a collection thirty-seven were never in — and a way
out of any of them without navigating there first.

The long press used to open the offline question by itself. That
question is one item in this menu now: there is one hold per row, and
while it was spent on a single action nothing else the tree can do had
a touch route at all. Nothing is lost — the tray on the row keeps its
tap, and the question gains a full-width control in place of a 30px
icon in a row shorter than the touch minimum.

The row-press handler moves to `collections_ui` with the rest of what a
collection row does; it lived in `library_ui` only because it opened
that prompt.
2026-09-07 19:59:44 +02:00
dtourolle 577bbd82b0 Return the action the drag offered, not the one this row prefers
Slint negotiates a drag action between source and target, and the
runtime clamps whatever `can-drop` returns against the set the source
allowed: an action outside it becomes `none`. Both drop targets here
named a constant instead of echoing what was on offer, and each named
the wrong one for half its traffic.

A collection row takes two kinds of payload. Photographs come from a
DragArea allowing `copy`; a collection being nested comes from one
allowing `move`. The row asked for `copy` unconditionally, so images
filed correctly and every collection dropped on a collection was
refused — nesting by drag has never worked. The trash had the same
fault mirrored: it insisted on `move` while the grid's cells allow only
`copy`, so it refused every photograph dragged to it.

Neither failure had anything to see. A clamped action is delivered as a
refusal, which looks exactly like a target that declined on purpose, so
the drag simply did nothing and left no error to search for.

`decide_drop`'s Reparent branch was tested and passing throughout. It
tests the decision, not the negotiation, and nothing was reaching it.
2026-09-07 19:59:44 +02:00
dtourolle bac5801618 Ask which collections a selection is filed in, and how much of it
`collections_for_image` answers this for one photograph and has no
counts, which is enough to badge a cell and not enough to offer a
removal: with forty selected and three of them in "Iceland", a sheet
that says only "Iceland" invites the user to take all forty out of a
collection thirty-seven were never in. `membership_of` returns the
count alongside the name so the row can say "3 of 40".

Chunked over the image list rather than one `IN (...)`, because the
list is a selection and a select-all makes it as large as the library —
past SQLite's bound-parameter cap on exactly the gesture most likely to
produce it. Counts are summed across chunks, so the answer is the one
the unchunked query would have given.

Smart collections are excluded by construction: they have no member
rows, so there is nothing a removal could do.
2026-09-07 19:59:44 +02:00
dtourolle af89433aee Offer the lens profile as a tick box, since applying it silently reads as absent
The develop panel's Optics group is three manual sliders: distortion, chromatic
aberration and lens vignetting. The automatic correction was already there — the
file's EXIF lens is matched against the bundled Lensfun database on open and the
coefficients are fanned out to all three — but nothing in the interface said so
except a line of grey text under the camera reading "· corrected", and there was
no way to decline it. From the outside that is indistinguishable from the
feature not existing, which is how it was read.

`dr-lens` states the rule this breaks: an automatic correction that silently
does nothing is worse than one the user can see is unavailable. The caption
satisfied the letter of it and not the point — a photographer looking for
"apply the lens profile" found three sliders and no switch.

So the profile is now a control. It is a capability rather than a flag on the
session, because everything a photographer sets travels one road: the capability
list feeds the generated panel, `Preset` captures it, the sidecar stores it and
the undo stack replays it. A bool on the side would have needed adding to each
of those four by hand and would have been forgotten in at least one — which is
exactly how the mask stack came to be missing from the history.

It is on by default, which is what `switch_on` is for: the coefficients are a
measurement of the lens that took the photograph, so accepting them is neutral
and declining them is the edit. The sidecar therefore stores nothing for the
ordinary case and the correction still happens.

The switch appears only where a profile was matched. A tick box on a photograph
whose lens the database has never heard of would be a control that looks
available and does nothing, which is the failure the rule above names rather
than an instance of following it — those photographs are told "· no profile" in
words instead, and one whose box is unticked now says "· profile off", which is
a third fact and not either of the other two.

Two things had to be built underneath. `ParamKind::Bool` was in the core's
closed enum and mapped to a row kind here, and had no control behind it in
`adjust.slint`: a parameter declaring itself a switch was flattened into a row
that drew nothing at all. Nothing shipped had one until now, so the gap cost
nothing and was invisible. And `Check` self-toggled, which is right for a
settings page that owns its value and wrong for a panel row that is a view of
the edit graph — the click would have answered by replacing the binding with a
literal, and the next undo or pasted preset would have moved the value with the
tick left where the finger put it. It now takes `controlled`, and the generated
row uses it.

The manual sliders are unchanged and still trim whatever the profile leaves, so
switching it off is "correct this by hand" rather than "stop correcting".
2026-09-07 00:27:47 +02:00
dtourolle 2cd49d1cb7 Measure the rail by counting it, so it can scroll instead of clipping
With the adjustment groups in the rail, a short window silently lost them.
At a 1500x680 window the rail drew Photo, Compose, Local, Repair, the seam and
"All", and then stopped: Optics, Light, Colour, Effects and Detail were not
scrolled off, they were gone, with nothing on screen to say so. The develop
view's primary navigation, unreachable by any means.

The Flickable was put here to prevent exactly that, and it could not, because
its viewport asked the layout how tall it wanted to be while the layout was
already taking its height from the viewport. Slint settles that cycle by
handing back the height it was given, so `max(self.height, preferred-height)`
could never exceed `self.height` and there was never anything to scroll.

Counting breaks the cycle. Every entry here is a fixed height by construction —
a tool is `rail-entry-height`, a group is a touch target — so the content is
six plus four fifty-fours plus a gap plus a touch target for each group and
"All", which is exact rather than an estimate and depends on nothing that
depends on it. Four tools and five groups come to 498, against the 340 a
680-pixel window leaves at 2x, and the difference is now scrollable rather
than absent.

Found by shrinking the window with the groups forced into the rail. The
interaction itself is unverified: synthetic input does not reach a Slint
window on this desktop, and the tablet was disconnected, so what is confirmed
is the arithmetic and the clipping it explains, not the scrolling it should
restore.
2026-09-07 00:17:08 +02:00
dtourolle 3dc7c184ee Optimise for release only, since every edit pays for a dev build
Dependencies were built at `opt-level = 2` even in dev, because wgpu and
image decoding are slow without it. That is still true, and it is what this
gives up: a debug run of the app, and the decode- and GPU-heavy tests, are
slower than they were.

What it buys is that nothing has to be optimised before it can be compiled.
That cost was paid on every edit, in every worktree, whether or not anything
was ever run — and there are sixteen worktrees, each with its own target
directory and no shared cache, so it was paid sixteen times over.

`[profile.release]` is untouched: `lto = "thin"` and `codegen-units = 1`
still apply where the speed is actually wanted.

If one crate turns out to be the one that makes a test unbearable, raise
that crate alone rather than restoring the blanket rule; the manifest says
how.

Incremental compilation is now on as well, but that lives in
`.cargo/config.toml`, which is untracked and per-checkout — so it is a local
change on this machine, not part of this commit.
2026-09-06 20:17:53 +02:00
dtourolle 3f6dbce2aa Name a lone control after its operation, so three cannot all read "Amount"
Benchmarks / CPU and I/O (per commit) (push) Successful in 10m31s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 1h15m4s
Build and test / Layer separation (push) Successful in 37s
Traceability / Requirement traces (push) Failing after 1m28s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Build and test / Android (aarch64) (push) Successful in 1h3m13s
The Detail group ended with three consecutive sliders labelled "Amount" and
nothing to tell them apart. They are dehaze, clarity and texture: each declares
exactly one parameter, and `ops/README.md` tells an author to reach for the
`amount` kind first, so all three named it the same thing.

The panel withholds a heading from a group of one, and the reasoning it gives
is sound — a lone control names itself, and the group reset it loses costs
nothing because the slider already resets on double-click. But that argument
rests on the parameter being named after what it does. It holds for exposure,
contrast, vibrance, saturation and brilliance, whose single parameter shares
the operation's name, and it fails for the three whose parameter is called
after its kind rather than its subject.

So a lone parameter now takes its operation's label — the name the withheld
heading would have carried. For the five that already agreed, nothing changes.

Found on the tablet, and only there: every row was correct, every label
resolved, and the panel was still unusable. The test that guards it asserts
over the real chain rather than a fixture, because the fault was a property of
what is actually declared — a fixture would have had to be written to
reproduce it, and would then only have proved itself.
2026-09-06 19:01:54 +02:00
dtourolle 124b2d99c6 Take the first control that moves the picture, not the first one drawn
`showing_the_original_leaves_the_edit_exactly_as_it_was` set row zero to its
maximum and then asserted the photograph was modified. It was not, and the
test failed on its own premise rather than on the thing it exists to check.

`EditGraph::capabilities` puts the lens corrections at the head of the list,
matching where they sit in the shader. Those carry profile coefficients rather
than parameters, so with no profile loaded a slider on one is a control with
nothing behind it: the graph stays neutral and the premise assertion fires.
The test was written against a panel whose first row happened to be an
adjustment, and it stopped being one.

So it now walks the rows until it finds a control that actually changes the
edit, which is what it meant by "row zero" all along. Still addressed by
index, so it still names no operation, and it no longer depends on where in
the chain the first *adjustment* happens to sit.
2026-09-06 19:01:53 +02:00
dtourolle 3994caba12 Write down the develop gestures, since only their author knew them
FR-UI-4 says a gesture with no visible counterpart is a feature only its
author knows about, and the vocabulary the application actually publishes had
two sections in it — the library grid and people. Develop had none. Every one
of its gestures was documented in the comment beside the `TouchArea` that
implements it, which is where the previous sixteen were before this scanner
existed, and unreachable to anybody not reading the source.

Thirteen now carry tags: magnify by pinch or wheel, pan a magnified frame,
fit and 1:1, hold to see the original, sample a neutral, undo and redo, step
through the folder, reset one control, show or hide a mask layer, choose a
group of adjustments, and copy and paste the settings. Each has a pointer and
a touch route, so none of them is keyboard-only.

Three bindings were genuinely missing and are added here rather than merely
described. Ctrl+C and Ctrl+V for the settings clipboard, which the Settings
panel's own comment has claimed existed for as long as the panel has and
nothing bound; and [ and ] to step through the adjustment groups. The groups
are whatever the operation set declares itself to be about, so there are as
many as the pipeline has and no key can name one of them — stepping is the
binding that survives a node being added, and "everything" is part of the
cycle rather than a way out of it.

Two are described and not bound. Resetting a control from the keyboard and
toggling a mask layer from the keyboard both need a notion of which control
or layer has focus, and the generated panel has none — the rows are a model
the repeater rebuilds, and inventing a focus ring for them is a larger change
than a keyboard shortcut. Both are reachable by pointer and by finger, and
the tags say so rather than promising a key that is not there.
2026-09-06 19:01:52 +02:00
dtourolle 901f51e6c4 Point at something grey and let the pipeline work out the rest
FR-DEV-3 has asked for "white balance (temperature/tint, and picker)" since
it was written, and only the first half existed. `WidgetKind::WhitePoint` was
in the vocabulary and `develop::supported` answered false for it, so the node
degraded to two sliders — correct behaviour that had quietly become the only
behaviour. Sampling a neutral is the first move of the global tonal pass and
every colour judgement afterwards is measured against where the grey was put,
so guessing at two sliders until a wall stops looking green is the wrong way
round.

The awkward part is that a picker genuinely needs to know how far a hundred
units of temperature move red against blue, and that number is declared in
the node's own file. So the inversion lives in `dr_pipeline::neutral` rather
than in the interface: the canvas hands over a colour, the core finds the
operation that asked to be driven by a pixel and bisects its declared
response until the sample comes back grey. Nothing in `ui/` names white
balance, and nothing holds a second copy of a response that would be wrong
the first time somebody adjusted the range. A bisection rather than a
closed-form inverse because only monotonicity is part of the bargain — the
expression is free to become a table tomorrow.

The result is rounded to the precision the control is drawn at, which is not
cosmetic: unrounded, sampling something already neutral lands a
ten-thousandth off zero, and the photograph comes back modified with an undo
step for a correction of nothing.

On the panel side this needed one distinction the generated path was
missing. `is_on_canvas` was being read as "and so the panel draws nothing for
it", which is right for a crop — four edge fractions are not controls anyone
drags in a list — and wrong for an eyedropper, which *writes* temperature and
tint and leaves them exactly the controls a photographer reaches for next.
So a sampling widget keeps its sliders and puts the affordance that arms the
canvas in the group's heading, built like the reset beside it. One click, one
sample, one history step: `Edit::Action` never coalesces, and there is no
hover preview to fill the stack with temperatures nobody chose.

Declaring the presentation also groups temperature and tint under one undo
step, where they were two. That follows from what `Presentation` means and
reads correctly — white balance is one decision — but it is a change, and
worth saying so.
2026-09-06 19:01:52 +02:00
dtourolle 2584b9ecbc Hold one key to see the photograph before you touched it
FR-DEV-7 asks for the current edit against the unedited original and nothing
implemented it. What the develop view had was history navigation, which
*changes* the edit rather than previewing against it — so the only way to
look was to undo, look, and redo, and that puts two real steps on the stack
at exactly the moment a photographer suspects they have overcooked a frame
and is least sure of what they are doing.

Holding the "Before" button, or backslash, renders the graph with every
adjustment stripped and hands it straight back afterwards: the same
suspend-render-restore shape the crop overlay already uses to show an
uncropped frame and an export uses to suspend the zoom. Nothing is recorded,
no rows are re-synced, and the photograph is still modified when the key
comes up — the panel goes on describing the edit the photographer has,
because only the canvas is answering a question.

The framing deliberately stays on. A held comparison is a question about
tone and colour, and re-cropping the canvas under someone's thumb would move
the detail they are comparing; worse, the zoom is a rectangle of the *framed*
image, so dropping the crop at 4× would quietly show a different part of the
photograph rather than the same part unedited. What the crop took away is
already compared in Compose, which shows the whole frame.

Not a split screen: that halves the working image on the tablet this column
was sized for, and the comparison photographers describe making is a flick
back and forth rather than two pictures side by side. Press-and-hold is one
gesture on a finger and on a mouse, which is what FR-DEV-3b's mapping wants,
and it has no mode to be stranded in — the button reports both edges, so a
press the system cancels puts the original down too.
2026-09-06 19:01:51 +02:00
dtourolle 9b674a88d8 Let the photographer look at the pixels, and keep looking
Noise reduction and capture sharpening are judgements about individual
pixels, and at a fitted view several of the file's pixels are averaged into
each one on screen. The frame therefore looks cleaner and softer than it is,
the photographer corrects for a softness the display invented, and
over-sharpening is the documented result. Nothing in the develop view
reached 1:1 at all: the wheel and the pinch zoom by ratios, the double tap
dropped straight to fit, and the only readout was a percentage nobody was
aiming at.

So the double tap now does what FR-UI-4 always said it did — toggle fit and
1:1 — and the zoom readout, which used to be a dead "Fit" button on a fitted
photograph, becomes the way in when there is nothing to clear. Z does the
same from the keyboard, and the back gesture goes out through the same
toggle so putting the magnifier down really puts it down. 1:1 is computed
from the file's own resolution against the viewport rather than fixed at
some multiple, because that is the only version of it that answers the
question the two detail controls are asking.

The point being inspected and whether the magnifier is up are held beside
the session rather than in it, on the argument focus peaking already makes:
a session is one photograph and this is a way of looking at a folder of
them. Checking the same eye across forty portraits is the reason to reach
1:1 in the first place, and a magnification that reset with the session
would make that forty zooms and forty pans instead of forty keystrokes. It
stays a viewing state throughout — the view is kept out of `is_active`,
`output_size`, the sidecar and the export, and an export still suspends it —
so none of this reaches the file.
2026-09-06 19:01:50 +02:00
dtourolle 43652bd613 Say where the positives really come from, not where they were going to
`calibrate.rs` claimed its positive pairs came from user confirmations "then
burst siblings, since FR-CULL-5 already groups bursts", and repeated it beside
the pair floor: "the positives are bootstrapped from bursts and a handful of
early confirmations". Neither is true and neither ever has been. Nothing in
the workspace pushes a burst pair into a `Pairs`; nothing pushes any pair at
all outside this file's own tests. The sentence was written while both halves
of docs/faces.md §8.1 were being planned together, describing a source that
was going to exist, and it has read since as a description of what the code
does.

The distinction matters more here than in most comments, because this file is
the one place in the subsystem allowed to say what a similarity *means*. A
reader who believes the fit is drawing on bursts believes a young library is
gathering positives on its own, which is precisely the opposite of the state
FR-CULL-9 legislates for — a library with no fit, no valid calibration, and a
reference curve it must not present as a measurement of itself. The floor of
200 positive pairs looks arbitrary under the wrong story and obvious under the
right one: confirmations arrive one at a time, from a person.

So the comment now says what is here — a positive is a pair confirmed onto one
person, and there is no second source — and keeps the burst idea where it
belongs, as §8.1's proposal, with the two reasons it is not in the code: this
crate is handed cosines and cannot see a catalog, and the purity of a burst
pair is a thing to measure before it is a thing to trust.

No behaviour changes; the arithmetic is untouched.
2026-09-06 19:01:50 +02:00
dtourolle 1c5c55b4c9 Let the photographer say which frame the burst stands for
`choose_representative` has been in the catalog since the grouping landed,
with tests behind it and nothing calling it. So the frame a folded burst drew
was always the earliest one, and the only way to disagree was to open the
group and leave it open — which is to say there was no way to disagree at all,
because a burst that stays open is a burst that was never collapsed.

The earliest frame is the right default and it is deliberately not a
judgement: nothing here scores a photograph, and FR-CULL-5 names the failure
that rule avoids. But the whole point of a burst is that one of the twelve is
better than the other eleven, and the person who knows which is the one
looking at them.

So a ring on each frame of an open group, ticked on the one the group folds
to. It is drawn only while the burst is open, because that is the one moment
the alternatives are on screen to be compared — offering the choice on a
folded burst would be asking about frames it is hiding. Bottom right, opposite
the count in the other corner, clear of the flag and the collection badge and,
deliberately, of the trash target: a slip between the ring and the fifth star
sets a rating, which is the harmless direction for an ambiguous press.

The mark stays live on the frame that already wears it. A disabled TouchArea
would let the press fall through to the cell behind it, so tapping the one
ring that is ticked would have opened the photograph — and pressing it is a
thing the user may mean anyway: it records the choice the default was making
silently, which then survives a regroup that finds an earlier frame.

Choosing repaints the badges instead of reloading the window, which is what
separates it from folding a group up. Folding changes what the grid's query
returns; this changes only which cell wears the tick, and the tick has to
leave the frame that was carrying it, so the whole window is refilled in the
one statement `sync_badges` already runs.

The gesture is documented where FR-UI-4 requires it to be documented: in a
tagged comment beside the control, which is the only copy. The gesture book,
the gesture document and the requirements matrix are regenerated from the tree
alongside it.
2026-09-06 19:01:49 +02:00
dtourolle 5fa4c0772b Speak the sidecar format every other editor already reads
FR-CAT-13 asked for standard XMP and nothing in the tree parsed or wrote a
byte of it. `keywords.rs` mentioned `dc:subject` in a comment about what a
keyword's text is for, `dr-export`'s metadata module said "neither is read by
`dr-decode` today" about its own half, and `dr-preset-xmp` reads a different
file for a different requirement. So a library imported from Lightroom could
come in and never go back out: a one-way door, which is not a thing a
photographer walks their archive through.

`core/dr-xmp` reads and writes the properties the requirement names —
`dc:subject`, `lr:hierarchicalSubject`, `xmp:Rating`, `xmp:Label` and the IPTC
core fields — from whichever shape the file happens to use. A property may
arrive as an attribute or as an element, inside a Bag, a Seq, an Alt or no
container at all, because the specification is not what wrote the file; so one
collector takes whatever is in a property and the declared shape decides only
how many values survive. `xmp:Rating="-1"` is modelled as Adobe's rejection
rather than folded into zero stars, since DarkRoom keeps those on two axes and
the mapping belongs where both are visible.

Writing is a rewrite rather than a serialisation, and that is the whole design.
An XMP sidecar is a shared document: the file beside a raw carries somebody
else's `crs:` settings and comments and namespaces, and rendering our record
over it would be data loss on every photograph but the first. The rule is
stated once, in the crate documentation and in `PROPERTIES`: DarkRoom owns
exactly those properties, identified by namespace URI and never by prefix, and
nothing else in the document. Everything unowned is copied through byte for
byte. A `Description` left empty once our properties come out of it is
withdrawn, which is what keeps a rewrite idempotent instead of adding a husk to
the file on every save.

Precedence is settled conservatively, because a standard XMP carries no
revision and no device and there is nothing in it to order two edits by.
Keywords union, following the rule `dr_catalog::merge` already makes for
assignments; every other field is taken only where DarkRoom holds none,
following `Version::merge`'s judgement rule, and a genuine disagreement is
reported rather than resolved so a caller can offer the reload the requirement
asks for. What is deliberately left open — when a reload may happen without
asking — is written down in the module rather than picked silently.

No new dependency: quick-xml was already in the tree for WebDAV and for
Lightroom presets. Nothing above the crate calls it yet, and `outstanding.md`
now says so along with the two smaller gaps, GPS and the filename convention.
2026-09-06 19:01:48 +02:00
dtourolle 68ebf5d78b Let a mask start from a tone or a colour, not only a shape
Every local adjustment began from a shape: painted, drawn with a handle, or
found by a model. So the only way to hold back a sky was to draw a line near
where it ended, and the only way to warm skin was to paint round it — both of
which put the edit's edge where the photographer put a gesture rather than
where the picture changes. A gradient across a treeline halos, and an
adjustment traced round a face stops on the outline of a hand.

MaskSource grows two variants that select by what a pixel *is*. Luminance
carries two bounds on the perceptual tone scale plus a softness; Colour carries
an arc of hue, a range of chroma, and one softness for every edge of both. Five
floats and three, so they diff, sync and merge per field under FR-NC-9 exactly
as a gradient's geometry does — the property a stored raster has none of, and
the reason the model's coverage had to sit beside its source rather than inside
it.

The pixels are the shader's business and nowhere else's. `mask.wgsl` takes the
demosaiced source as a sixth binding and two new modes read it: decode, balance,
pull a clipped photosite back to neutral, apply the camera matrix, then weigh
the band. Nothing crosses to the CPU but the numbers and the matrix, and each
mask texel averages its own footprint in the source, so a band lands on the tone
an area is rather than on whichever texel a proxy grid happened to land on.

The photograph it measures is the one the camera recorded, before this edit. A
band over the edited result would slide out from under the edit as the edit was
made — raising the highlights would change which pixels counted as highlights,
and the slider would chase its own mask.

Feather, falloff and morphology stay off a range layer, which is what
`shapeable` already meant. All three are functions of the signed distance from
a boundary, and a range has no boundary to be at a distance from; its edge is
the softness of its own band, in the band's units. Offering them would be four
controls that move and change nothing.
2026-09-06 19:01:48 +02:00
dtourolle 81b1ae8c42 Measure the haze from the picture, and divide it back out
Four files named dehaze as a member of the compositional detail family —
`detail.rs` twice, `dr-gpu`'s detail module, `ops/README.md` and
`capture_sharpen.rs` — and no such node existed. Every one of them was
describing the family by listing clarity, texture and a control the
photographer could not reach.

Haze is the one degradation the controls already in the chain cannot
remove, and the reason is spatial rather than tonal. Scattering
composites an airlight over the scene in proportion to distance, so the
lift is per-pixel: a black point that clears the mountains crushes the
foreground, and a contrast curve that clears the mountains does the
same. So the node has to estimate the transmission at every pixel, which
is the dark-channel prior — the local minimum over the channels and over
a patch is the airlight that has been added there — and then invert the
scattering model with it.

The airlight is taken as neutral and as unit, which removes the one part
of the published method this stage cannot perform. Estimating it
properly is a whole-frame reduction, and the detail chain has none: it
hands each pass the pass before it. It is also unnecessary, because
white balance is the first node in the chain and has already driven the
illuminant to grey, so only the magnitude is unknown — and an unknown
magnitude on the veil is a scale factor on the amount slider, which the
photographer is setting by eye regardless.

The patch is a fraction of the frame's shorter edge, through
`RenderScale::frame_fraction`, and never a count of pixels. It has to be
wide enough to contain something dark and narrow enough that what it
measures is still local, and both of those are statements about how much
of the composition it covers — so it must cover the same proportion of
the picture on a proxy as in the export, or the file is sharpened for a
patch three times narrower than the one that was tuned on screen.

Affording it needs an identity a Gaussian does not have. Erosions
compose by adding their structuring elements, so the minimum over a run
of d followed by the minimum over k points spaced d apart is the exact
minimum over the whole kd window. At the square root that is 16 taps
rather than 61 at 4K, and it is the same filter rather than an
approximation of one — which is the difference from the strided kernel
`local_contrast` refuses, where sampling an image that is not
band-limited aliases into the base and comes back as mottling.

It runs first among the compositional detail nodes, at order 125: after
noise reduction, because dividing by a transmission below one amplifies
the noise in the veiled distance by exactly the factor it recovers the
contrast by, and before clarity and texture, coarse before fine, so that
their base is computed on the picture the veil has left rather than on a
modelling about to be divided out.

What it cannot honour is the placement dehaze most wants. It shifts
colour — it subtracts a grey term and rescales, so saturation changes
wherever the veil is thick — and the colour work would ideally be
correcting the picture that leaves here. The detail stage runs as a
group after every point operation, because a neighbourhood pass is a
separate dispatch over a texture the fused pass has finished writing, so
an order placing this node ahead of `vibrance` would be a lie the chain
cannot tell. Interleaving would mean splitting the fused pass in half
around it, at the cost of a second full-frame dispatch and intermediate
for every edit in the catalogue whether it dehazes or not. The
declaration records that rather than leaving it to be rediscovered.

FR-DEV-18 is added to the requirements register alongside it. The tag
had nowhere to point, and an orphan tag fails the traceability gate
rather than quietly counting for nothing.
2026-09-06 19:01:47 +02:00
dtourolle 7c3e1d2c54 Let the shadows and the highlights carry a colour the picture never had
The colour mixer is the only chromatic control in the chain, and it can only
turn a hue that is already in the frame. Ask it for cool shadows against warm
highlights and it has nothing to take hold of: the shadows of a correctly
balanced photograph are near enough neutral that there is no band there to
turn, and a monochrome conversion hands it a picture with no hue in it at all.
Split toning is the oldest look in the book and every developer worth comparing
against ships it; there was no way to reach it from here.

So colour_grading, declared like any other node — a hue and a strength for the
shadows, the midtones and the highlights, and a global cast over the frame. It
targets a tonal range rather than a hue, which is the whole difference between
the two controls: it puts colour where none was rather than turning what it
finds. It sits at 105, after the mixer has had the last word on the colours
that are in the picture and before the detail stage.

The mechanism is one helper. Three cosines 120 degrees apart are the hue wheel
written directly as an RGB direction, and their sum is zero at every angle, so
exp2 turns them into three gains whose product is exactly one — a cast tilts
the balance without moving the level. A grade that doubled as an exposure
change is the failure that has the photographer chasing brightness with a
colour slider, and it is corrected with a control that cannot reach it. The
three tonal weights partition the scale rather than overlapping, the midtones
being whatever the two ends leave, so setting all three to one hue is exactly
the global cast and a split tone does not colour its own midtones as a side
effect of its halves meeting. Full strength is half a stop on the leading
channel, the ceiling white balance already holds itself to.

Neutral is declared rather than inferred, which is what `active:` is for. A hue
with no strength behind it is a direction with no distance, so under the
default rule nudging one would have put the node into every fused shader for a
change nobody can see. Summing the strengths is zero exactly when all four are,
and they cannot go negative to cancel each other. The opposite reading —
neutral as "nothing has been touched" — fails the other way round: red is hue
zero, so a grade toward red never moves a hue off its default and would never
have been applied at all.

It asks for a colour wheel, the widget the descriptor vocabulary has been
carrying with no operation behind it. Nothing draws one yet, and that is fine
by construction: the panel takes the first widget it implements and falls
through to sliders otherwise, so this arrives as eight ordinary controls that
work. Each parameter is named for its own range for exactly that reason — in a
flat list, four sliders called "Hue" are four controls nobody can tell apart.

FR-DEV-12 is written into requirements.md beside it. A TRACES tag naming a
requirement that is not defined there is an orphan, and the traceability gate
fails on those rather than quietly counting them. The label catalogue gets one
line for the operation's display name; the eight parameters derive correctly
and are left to.
2026-09-06 18:51:40 +02:00
dtourolle efa9d84aad Correct the lens first and settle the grain last
`Attribute::ALL` has claimed since it was written to be roughly the order a
photographer works in, and 474dcf0 moved `Compose` to the front on exactly that
argument. Both ends went on contradicting it.

`Optics` sat fifth, so the column offered the lens corrections after the tones
they change. Removing a vignette brightens the frame; an exposure judged before
that correction has to be judged again after it, which is the definition of the
wrong order. `Detail` sat fourth, so sharpening and noise reduction — the only
work here that depends on everything above it, and the only work that cannot be
judged at fit view at all — were offered before the lens had even been put
right.

So `Optics, Compose, Tone, Colour, Effect, Detail`. `Optics` leads even
`Compose` because it is not a decision about the photograph at all: it is
undoing what the equipment did, a property of the capture rather than a choice.
`Effect` after `Colour` is a look laid over a settled picture, and is the one
slot that is genuinely arguable — a spectral film simulation declares `renders`
and replaces the base curve, which is a case for treating it as foundational
instead, and e235e99 filed film under `Effect` only days ago. The doc comment
records that tension rather than pretending to settle it; an array of six
cannot say "last, except when it is first".

`decl::Attr::ALL` moves with it. It is the second spelling of one vocabulary,
compiled by `build.rs` where `descriptor` is not visible, and the agreement
test in `declared/mod.rs` zips the two positionally — that test is what caught
the last reorder, and it would have caught this one.

Nothing persists a position in this list, which is what makes the reorder safe
rather than merely tidy. `Scope` packs one bit per attribute indexed by
`Attribute::ALL`, but `bits` is private, has no accessor and no `serde`; what
reaches a settings file is `develop.copy_attributes`, a list of names read back
through `Attribute::from_name`. A photographer's copy scope survives untouched
— only the order the names happen to be written in changes.

6a97fdf is why this is worth a commit now rather than a shrug: the
contradiction was harmless while the list only fed a row of chips nobody reads
in order, and stopped being harmless when the same list began driving a column
read top to bottom.

The matrix follows the two files' shifted line numbers.
2026-09-05 21:17:56 +02:00
dtourolleandClaude Opus 5 59917c5183 Call the tool Compose, since that is what its panel says
Benchmarks / CPU and I/O (per commit) (push) Successful in 14m35s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 1h20m20s
Build and test / Layer separation (push) Successful in 51s
Traceability / Requirement traces (push) Successful in 2m8s
🐳 Android image / Build and push (push) Successful in 5s
Build and test / android-image (push) Successful in 5s
Build and test / Android (aarch64) (push) Successful in 1h5m51s
The rail entry read "Crop" while the panel it opens is headed COMPOSE and the
button leaving it said "Done Cropping". One mode, three names, and the odd one
out was named after a single control rather than after the decision — which is
what made cropping look like a category of its own in the first place.

Straightening, the quarter turns and the flips are already in that panel, and
perspective will be. `ViewMode.crop` keeps its name: it identifies a canvas
interaction, which is exactly what it still is.

Found by looking at the running application rather than by reading, which is
also how the two halves of this were noticed to disagree at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 18:29:08 +02:00
dtourolleandClaude Opus 5 e235e99cce Move the film to Effect in the descriptor that is actually read
An earlier commit claimed to move `film_sim` from `[tone, colour]` to
`[effect]` and did not. It edited `ops/film_sim.yaml`, where `attributes:` is
read, validated against the vocabulary, and then dropped: a `rust:` node
publishes its own descriptor, and the type still said tone and colour. The
stock went on appearing in the Light group beside exposure and again in Colour
beside white balance, exactly as before, and every test passed.

Nothing caught it because nothing could. The declaration parsed, the parity
tests compare ids rather than attributes, and an operation filed under the
wrong groups renders perfectly. It surfaced only on screen, as a missing
Effects tab — which is indistinguishable from a category that genuinely has
nothing in it, and is precisely how `Optics` looked for as long as it was
empty.

So three changes rather than one:

`FilmSim`'s descriptor declares `Attribute::Effect`, which is the move the
earlier commit described.

`attributes:` joins the keys a `rust:` node may not carry, beside `params`,
`uniforms`, `wgsl`, `helpers`, `define` and `label`. The rule was already
written — "its descriptor comes from the type" — and attributes were the one
field that slipped past it. A key that is silently ignored is worse than one
that is rejected, because it reads as though it worked; the eight hand-written
declarations lose a line that never did anything.

And a test asserts that every attribute the chain carries reaches the tab
strip. That is the property that was actually broken, and its failure mode is
invisible from every direction: the controls exist, they are in the shader,
and there is no way to filter to them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 18:28:59 +02:00
dtourolleandClaude Opus 5 6a97fdf6f9 Put the adjustment groups in the rail where a finger is driving
Reported from the tablet: the tool rail is very useful there, and the same
interface under a mouse and keyboard is not. That is `ui-navigation.md` D-N2's
central assumption failing in use, and the interesting part is which half of
it failed.

D-N2 was right that platform is the wrong axis and width is the wrong axis: a
tablet in landscape wants what a desktop wants, and a desktop window dragged
narrow wants what a small screen wants. `apply_layout_class` still decides the
layout class from the window and nothing here changes that. What D-N2 got
wrong is the sentence "touch changes hit regions, not layout" — it identified
input as the real difference between the targets and then assumed that
difference could never reach the layout.

Two controls answer one question — which group of adjustments am I looking at
— and neither is better in general. A horizontal strip above the column is one
gesture to a target the eye has already found, and it pans when the operation
set is rich, so a group can sit off the end with nothing saying so: a pointer
user tolerates that, a finger user never discovers it. The same list down the
rail is every entry visible at once, each finger-sized, on the edge of the
screen the hand is already holding, and it costs no width because the rail is
already there.

So `ToolRail` grows a second section, and `GroupStrip` stands down when it
does. The two are never both on screen, which is why they can share
`adjust-tab-picked`: Rust is not told which was pressed and has no reason to
want to. Mode and group stay independent axes as N1 requires — one entry lit
in each section, and choosing a group while a tool is held still filters
without putting the tool down.

They stay drawn differently, which N1 also required. The tools fill with
`active-dim` and invert their ink; the groups take a bar down the leading
edge — the strip's underline turned ninety degrees — so a lit entry says which
kind of state it is without the reader having to remember which section it was
in. The rule between the sections is the second signal.

The rail scrolls now. Its own note argued against a Flickable because "this
list is four entries written in this file"; with the groups in it the list
comes from the operation set, which is exactly the "something the user's data
decides" that note excluded this control from.

The axis is input, and it is a preference because the automatic answer is a
guess that cannot be made reliable. Neither platform can be asked what the
user is holding: an Android tablet in a keyboard case is being driven like a
desktop, and a touchscreen laptop is whichever its owner says.
`dr_plat::is_touch_first` reports the usual case per platform, and
`GroupNavigation` lets it be overridden. Settings names what Automatic
resolves to on this device rather than leaving it to be found by pressing.

D-N6 records the reversal beside the decision it reverses, including the half
that still stands and the question it opens: whether Local is a mode at all,
or a scope that would collapse the two sections into one list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 18:14:03 +02:00
dtourolleandClaude Opus 5 474dcf0bf6 Let Compose lead the attributes, as their own doc always said
`Attribute::ALL` claims to be "roughly the order a photographer works in" and
then listed framing fifth, behind tone, colour and detail. Framing is the
first decision made about a photograph and the one every later judgement is
made inside — there is no sense balancing tones across a frame about to lose a
third of its width.

The contradiction was harmless while the list only fed a row of chips nobody
reads in order. It stops being harmless now that the same list drives a column
read top to bottom.

`declared::Attr::ALL` moves with it. The two are separate spellings of one
vocabulary and a test asserts they agree, which is what caught this rather
than the order silently disagreeing between the YAML front end and the crate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 18:13:44 +02:00
dtourolleandClaude Opus 5 1ad35e2b87 Read the library's sidecars, so a cull done elsewhere arrives
Judgements only ever travelled outward. A rating went to the catalog and to the
photograph's sidecar, the sidecar reached the server, and there it stopped: the
scan indexes files, `derived_sync` exchanges thumbnails, face shards and
collections, `dr_catalog::merge` reconciles everything in a catalog except
`versions.rating` and `versions.flag`, and the one sidecar reader that existed
ran when a single photograph was opened in develop and handed its answer to the
develop graph. `JobKind::ReadSidecar` was declared for exactly this when the job
queue was written and was never enqueued or handled anywhere.

The grid draws `versions.rating`. So a day of culling on the tablet could not
reach the laptop by any path the application had, and the laptop's catalog says
so plainly: 23,568 images, one of them judged.

`pull_sidecars` closes it, off the back of work the scan already does.
`dr_sync::scan` reports the `.drsc` files it meets in listings it was making
anyway — no extra request, and a directory whose ETag is unchanged is still
pruned before it is listed at all. A new `sidecars` table records the ETag of
each one this device has taken in, so the fetch is one GET per sidecar that
genuinely changed rather than one per photograph. A library nobody has edited
costs nothing.

The judgement is taken rather than maximised. The sidecar is the authoritative
store and the fuse has already settled any contest between devices on
`revision`, so lowering a rating from four to one on the tablet lowers it here —
taking the larger would have refused every demotion the photographer ever made,
which is most of what a second pass over a shoot is. A zero is the exception: it
means *never judged*, not "judged zero", so a sidecar carrying none cannot erase
a star this device holds. That is `merge_judgement`'s asymmetry and it carries
the same known cost — clearing a rating does not propagate.

A sidecar names a stem, so both halves of a RAW-and-JPEG pair are judged: they
are one photograph (FR-CAT-11) sharing one document, and judging only one of
them would leave the grid disagreeing with itself over which it drew. The `LIKE`
that finds them is a filter, not the decision — `sidecar_path` is applied to
every candidate, because a folder is entitled to contain a `%` and a rating
landing on the wrong frame would be silent and permanent.

Failing to read one is not a failure to scan: the ETag goes unrecorded, the
ratings already here stay where they are, and the next scan tries again. The
count is reported to the status line as well as the log, because a grid that
silently gains three hundred stars is indistinguishable from one that has gone
wrong — and because while this number was structurally zero there was nothing to
tell the photographer their cull had not arrived.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 16:14:03 +02:00
dtourolleandClaude Opus 5 ca2a135e28 Let two devices name the same photograph's version the same way
A version's uuid is the identity a cross-device merge keys on, and it was
minted at random, per catalog, per image. Two devices indexing one Nextcloud
library therefore held two different uuids for the same photograph — so the
sidecar they shared collected a `default = 1` block each, `Version::merge` was
never handed a matching pair to reconcile, and an afternoon's culling on the
tablet did not exist as far as the laptop was concerned.

`crate::merge` has said so in a comment since it was written: version uuids do
not reconcile across devices, a uuid-keyed join unions nothing, so keywords are
landed on the local default version instead. It named the problem and worked
around it. `rating`'s own comment asserted the opposite — that generating the
uuid here was what made it a cross-device identity — and `library::amend`
repeated the claim. Uniqueness was never the difficulty; agreement was.

`derived_version_uuid` computes it from `oc:fileid` instead. The server assigns
that integer, every client pointed at the library sees the same one, and it
survives a server-side rename and move — the three properties that already made
`ASSIGN_BY_FILE_ID` prefer it to a content hash. The layout is a UUIDv8 (RFC
9562, an application-defined form) carrying all sixty-four bits verbatim across
the variable fields with a fixed tag in the node field, so the mapping is
injective by construction rather than by a hash's good behaviour, and a uuid in
a sidecar can be read back to the file it belongs to by eye.

A library with no server behind it has no shared identity to derive and keeps a
generated one. The split is still reachable there if the folder is synced by
something else; `Sidecar::fuse_default_versions` repairs that case rather than
preventing it.

Deriving it for new rows alone would have fixed nothing — every image in an
existing library already has a version, so every one of them would have carried
on writing to its own rival identity. `align_default_version_uuids` moves them,
and runs from `schema::backfill` on every catalog open. It selects on the tag
in SQL, so a catalog already realigned matches no rows and writes nothing, and
it declines rather than fails where a virtual copy already holds the target.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 16:13:58 +02:00
dtourolleandClaude Opus 5 601c984894 Fold a photograph's rival default versions back into one
Picking the newer of two default versions stopped the wrong edit being shown,
but it did not close the split: the losing version stayed in the file, and a
device holding disjoint work — a crop made here, an exposure change made there
— still contributed only one of the two.

Worse, the next write made it larger. `amend` looks its version up by uuid,
neither of the two was ours, so the miss minted a *third* `default = 1` block
and the file grew one rival per device per photograph.

`Sidecar::fuse_default_versions` folds them down. The version with the highest
`(revision, modified)` is the accumulator and every other default is merged
into it as the remote, which is what makes the fold order-independent —
`Version::merge` raises its own revision to `max + 1` as it goes, so merging a
chain in ascending order stops being ascending after the first step and a third
device would be dropped. Contested values resolve to the winner, disjoint keys
survive from both sides because the merge is key-wise, and ratings come across
under `merge_judgement`, so a device that never judged the frame cannot erase
one that did.

The result is a function of the file's bytes alone, so two devices that fuse
independently reach the same document and converge instead of overwriting each
other.

Called wherever a sidecar is parsed:

- `amend`, with the write's own uuid, so the fold lands on the identity this
  device is about to use and the lookup below it hits instead of missing.
- `spawn_sidecar_fetch`, so opening a photograph shows everything done to it
  rather than whichever half won.
- `drain_one`, because `merge_into` reconciles by uuid and would otherwise
  publish the split rather than resolve it.
- `presets::load_local` and `save_local` — a local sidecar's folder may be
  synced by something else entirely, and gets the same split.

A file with one default under the expected uuid comes back byte-identical, so
this costs nothing on the ordinary write and no sidecar is uploaded merely for
having been read.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 16:13:53 +02:00
dtourolleandClaude Opus 5 98fcf8e98e Ask which edit is newer, not which uuid sorts first
A photograph edited on two devices ends up with two `[version]` blocks in one
sidecar, both marked `default = 1`. Five of the thirty-eight sidecars in the
local cache are in that state right now.

`default_version` answered with the first `is_default` it met in map order,
and the map is keyed on uuid — so which device's work the photographer saw was
decided by which randomly minted uuid happened to sort lower. On
`IMG_20130625_0033` that is `0545c20a` over `679fe872`: a four-star rating from
the twenty-first of August standing in front of the one-star made on the
thirtieth, with nothing anywhere saying the newer judgement existed.

Resolved by `(revision, modified)` instead, which is the discriminator
`Version::merge` already uses — revision first so that a device with a skewed
clock cannot win by claiming a later timestamp (FR-NC-8), and the timestamp
only to break an exact tie.

This makes the reader pick the right one. It does not make the two converge:
the edit that lost is still in the file, and a device that holds disjoint work
— a crop here, an exposure change there — still only contributes one of them.
That is the next commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 16:13:49 +02:00
dtourolleandClaude Opus 5 a56baa9042 Let rustfmt have the assertion it reflowed
The closure taken down to `&dyn Operation` needed its call site re-wrapped,
and I wrapped it by hand rather than letting rustfmt decide: it fits on one
line at the workspace width. `cargo clippy` was run on the change and
`cargo fmt --check` was not, which is the whole of how it got through — the
two catch different things and CI runs both.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 15:37:46 +02:00
dtourolle f0ee53ec09 Merge: the optical corrections, connected at last
Four files' worth of lens correction existed, was tested, and had never
touched a photograph. `compose_warps` had no callers, `dr-lens` had no
dependents, and `ops/vignetting.rs` had no declaration in `ops/` — so
`Attribute::Optics` was a category the tab strip could only ever filter out
for having no rows in it.

Distortion, chromatic aberration and lens vignetting now render, carry their
parameters through the sidecar and the undo stack, and take their coefficients
from the Lensfun database when the file names a lens it knows. The panel says
which of "no lens recorded" and "no profile for this lens" it is, because an
automatic correction that silently did nothing is worse than one visibly
unavailable.

One real bug on the way: `lens.rs` and `framing.rs` both documented the
shader's `p` as corner-normalised, and it is not — its length at the corner is
`0.5 * length(aspect)`, about 0.901 on a 3:2 frame. Every Lensfun polynomial
would have been evaluated short of where it was fitted, by a factor varying
with the aspect ratio, which reads as a correction that is merely too weak.

The category vocabulary moved with it. `Attribute::Geometry` is `Compose` —
named for the photographer's decision rather than for the maths it shares with
the lens corrections — and a film stock stopped claiming to be both tone and
colour, which had put "Kodachrome" in two groups it belongs to neither of.
2026-09-05 15:35:44 +02:00
dtourolleandClaude Opus 5 2841eaf9a1 Take the layer-chain test's closure down to &dyn Operation
`clippy::borrowed_box` is denied by the workspace lint set, and the closure
added with the optics exclusion took `&Box<dyn Operation>` — a borrow of the
box rather than of the thing in it, which says nothing the plain trait object
does not.

Caught by `cargo clippy --workspace --all-targets -- -D warnings`, which is
what CI runs and what the workspace tests do not: a lint on test code only
appears when the tests are compiled as a clippy target.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 15:14:15 +02:00
dtourolleandClaude Opus 5 c4ddcbe0f7 Look the lens up and say plainly whether one was found
`dr-lens` has held a complete Lensfun lookup — distortion, TCA and vignetting
coefficients from a lens name, a focal length and an aperture — with no
dependents anywhere in the workspace. The three corrections it feeds now
exist in the graph, so this connects the two and finishes the chain.

The coefficient structs stay duplicated. `dr-pipeline` is organised around
having no dependencies so its codegen is testable without a device or a
database (ARCH §6.5a), and `dr-lens` carries an XML parser and 5.5 MB of
profile data. Neither crate can convert to the other, so the conversion goes
above both, in `develop.rs`, which is the only place that sees them together.

Both traits grow the same defaulted door. The optical corrections do not sit
on the same side of the fetch — distortion and CA rewrite coordinates and are
`Warp`s, vignetting applies a gain to the pixel already there and is an
ordinary node — and fanning a profile out by which trait each happens to
implement would make the caller reason about that distinction. Each correction
takes its own share of the whole profile instead, and `set_lens_profile` walks
both lists identically.

The lookup happens in `set_source_metadata` rather than in its caller, because
that is the one place a session is told which file it came from. Doing it
there makes it unforgettable, in the shape `FilmRebake` already uses for the
other derived thing — and, more to the point, makes *clearing* unforgettable:
a session that opened a second photograph while still holding the first one's
profile would correct it for the wrong optics, invisibly, in a way that looks
exactly like the lens.

It needs the whole shot and not just a name. Distortion is interpolated across
a zoom's focal range and vignetting depends strongly on aperture — a fast
prime can be two stops down in the corners wide open and clean by f/8 — so a
lookup missing either returns coefficients measured for a shot nobody took.
Missing any of the three refuses rather than guesses.

A profile is derived, not persisted: it comes from the file's EXIF and a
database, so it is not a parameter, not in the sidecar and not undoable. What
is an edit is the manual trim beside it, which each correction composes with
the measurement — so a photographer can lean on it, override it, or work
without one.

`InfoPanel` gains a lens line, and it distinguishes three cases rather than
two. `dr-lens` states the rule it exists for: an automatic correction that
silently did nothing is worse than one the user can see is unavailable. A
session with no header draws nothing, a header naming no lens reads "Lens not
recorded", and a lens the database has never heard of reads "· no profile".
Collapsing the last two would send somebody hunting for a profile that was
never missing — which, for third-party and adapted glass, is the ordinary case.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 15:11:32 +02:00
dtourolleandClaude Opus 5 a1165ef182 Put the coordinate-domain lens corrections into the graph
`lens.rs` has held a `Warp` trait, a composer and two implementations —
distortion and lateral chromatic aberration — since they were written, and
`compose_warps` was called by nothing outside its own tests. The corrections
existed, were correct, and never touched a photograph.

`EditGraph` now holds them, and `compose_full` emits them between the framing
prologue and the fetch. Distortion first, then CA: each warp receives the
position the previous one produced, and lateral CA is a magnification about
the optical axis of the *undistorted* frame, so measured on a barrel-distorted
one it would be fitted to a radius no profile describes.

They reach the panel the way framing already does — through `capabilities`.
That was the one open question and existing practice answered it: framing is
also not an `Operation`, also has parameters a photographer sets, and also
arrives through that list. Because `Preset::capture` walks the same list, the
sidecar, the clipboard and the undo stack carry a warp's parameters with
nothing registered anywhere, and no file under `ui/` names one (FR-DEV-3a).

`state()` destructures `EditGraph` field by field precisely so that a new
field cannot be forgotten, and it was not.

Chromatic aberration is the only thing that samples per channel, and
`splits_channels` is what keeps everything else from paying for it. Red and
blue are fetched from positions green is not — green is the reference and
never moves, so a wrong correction still leaves one channel sharp rather than
softening all three. With no CA in the chain the single-fetch path is emitted
instead.

The interpolating sampler is now chosen by framing *or* an active warp. Asking
framing alone would have nearest-neighboured a distortion correction on an
unstraightened frame, and that aliasing reads as a bad profile rather than as
a missing filter.

The warps go in the geometry invalidation key rather than the colour one: they
decide which source pixel a colour is read from, so a tile cached across a
distortion change would keep drawing the previous correction. The pipeline
cache needs nothing new — `hash_source` already covers the generated body, and
uniform values never enter it, so arming a warp recompiles and dragging it
does not. Both are asserted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 15:03:16 +02:00
dtourolleandClaude Opus 5 e17b909d41 Connect the lens vignetting correction to the pipeline
`ops/vignetting.rs` has carried a complete descriptor, polynomial, helper
and test suite without an entry in `ops/`, so it was never in `chain()`.
It reached no photograph and no panel, and `Attribute::Optics` was an empty
category in consequence — filtered out of the tab strip for having no rows,
by a chain that had never been given its only member.

Declaring it needs the one thing the operation was written against and
which did not exist. `wgsl_body` reads `radius`, and the module claimed
"the composer publishes `radius` in the shader prologue for exactly this
reason". It did not. `sample_source` now does, in both sampling branches,
beside the `source_px` it already published for the same class of caller.

It is corner-normalised there, which is the part that is easy to leave out.
`p` spans ±0.5·aspect, so its length at the corner is 0.5·length(aspect) —
about 0.901 on a 3:2 frame, not 1. Lensfun's polynomials are fitted against
a corner radius of 1, so passing `length(p)` straight in evaluates every one
of them short of where it was measured, by a factor that changes with the
aspect ratio. It would have read as a correction that is simply too weak,
which is indistinguishable from a bad profile. Both `lens.rs` and
`framing.rs` asserted the normalisation `p` does not have; corrected.

`order: 5` puts the correction ahead of the tonal stages, and the ordering
is load-bearing rather than tidy. Recovering a corner means dividing by an
attenuation below one — about two stops for a fast prime wide open — so run
after the highlights have been rolled off and clipped, the lift has nowhere
to go and the corners posterise instead of brightening.

`layer_chain` now drops `Optics` as well as the neighbourhood operations.
A local vignetting slider would have worked, which is what makes it worth
excluding: `radius` measures from the centre of the whole photograph and a
mask cannot move the optical axis, so it would lay a frame-centred radial
ramp across the picture and multiply it by the mask. The existing exclusion
covers operations that move and do nothing; this one covers an operation
that moves and does something its name does not promise. The rule both
share is now written down.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 14:34:24 +02:00
dtourolleandClaude Opus 5 8e7b1350bf Name the frame's category for the decision, not the maths
`Attribute::Geometry` becomes `Attribute::Compose`, and `film_sim` moves
from `[tone, colour]` to `[effect]`.

Two categories were doing the wrong job. "Geometry" describes what crop,
straighten and the quarter turns do to coordinates — but it describes lens
distortion correction exactly as well, and that is not a compositional
choice at all. Naming the attribute for the photographer's decision is what
separates it from `Optics`: one is what the lens did, the other is what
they chose. The maths the two have in common is not the thing worth
filing them under.

A film stock declared both `tone` and `colour`, so "Kodachrome" appeared in
the Light group beside exposure and again in Colour beside white balance —
two places, neither of which is where anyone looks for it. It is neither:
`Effect` is defined in this same file as "applied rather than corrected — a
look, not a fix", which is what a stock is. That it moves tone and colour
is true of every look, and is not what the attribute is for.

`from_name` still accepts "geometry" on the way in. That string is
persisted in `develop.copy_attributes`, and an entry it fails to parse is
not an error — `presets::scope_for` logs it and drops it — so without the
alias an existing settings file would have quietly narrowed what a paste
carries. `name` writes the current spelling, so the file migrates itself
the first time it is saved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-05 14:34:03 +02:00
dtourolleandClaude Opus 5 95e854b5a2 Regenerate the gesture vocabulary over the library work
Traceability / Requirement traces (push) Successful in 1m57s
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m42s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / android-image (push) Successful in 5s
🐳 Android image / Build and push (push) Successful in 5s
Build and test / Android (aarch64) (push) Successful in 1h5m40s
Build and test / Layer separation (push) Successful in 48s
Build and test / Desktop (Linux) (push) Failing after 1h12m34s
`Traceability` stayed red after the matrix was regenerated, on its other
gate: `gestures-check`. Same cause, second artefact. `docs/gestures.md`
cites each gesture by `file:LINE`, and the library UI work moved the two
selection-mode gestures down a hundred lines — 3923 -> 4032 and
3940 -> 4049 in `ui/dr-ui/ui/library.slint`. Nothing about the gestures
themselves changed.

`ui/dr-ui/src/gesture_book.rs` was already current, so this is the doc
alone: 70 files scanned, 16 gestures, 2 places, gate PASS.

Worth knowing for next time: `tools/ci-local.sh traceability` runs the
self-test, the coverage gate and the matrix, but not `gestures-check`, so a
clean local run does not prove this workflow green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 03:06:31 +02:00
dtourolleandClaude Opus 5 c1069ce07e Release 0.10.1
Benchmarks / CPU and I/O (per commit) (push) Successful in 13m30s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / android-image (push) Successful in 3s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / Layer separation (push) Successful in 46s
Traceability / Requirement traces (push) Failing after 1m36s
Build and test / Desktop (Linux) (push) Failing after 24h17m53s
Build and test / Android (aarch64) (push) Canceled after 0s
A bug fix release. Nothing here changes what DarkRoom is for since 0.10.0;
it changes how often it does the thing it already claimed to do.

The edits that went missing. A subject or category layer was stored as
identity alone, on the reasoning that the pixels were reproducible by
re-running the model — true, but nothing re-runs one except a photographer
pressing "find subjects". So a reopened photograph rendered without its
local adjustments and then saved that state back, and a batch export wrote
three hundred files without the edits their photographer had made, over a
log warning. The coverage now travels in the sidecar. Export also learned
to write the photograph rather than the canvas, and to carry the
photograph's header when the export is made from develop.

Face indexing was reading previews. It ran against a proxy and then
recorded the result as though it had seen the photograph, so faces smaller
than the proxy could resolve were not missed, they were *concluded absent*.
Indexing now runs on the native render, runs made against proxies too small
to find a face are forgotten rather than trusted, and a sweep that fails
everything says so instead of reporting a clean pass.

Where you were. The photo roll opens on the frame it opened with, develop
returns you to the photograph you were editing, the grid keeps its place
when another screen covers it, and the photographer's position now travels
between devices rather than being rediscovered on each.

Startup. The catalog opens on a worker and the bundled models unpack on
one, so a launch is no longer a page-by-page read on the way to the first
frame; the app says it is starting before there is anything to say it with.
The People rail builds the rows you can see, keeps portraits off the
blocking path, and withholds the empty groups that used to fill it.

Segmentation gained the half it was missing: the colour gate decided what
belonged to a category and had no way to decide where its edge fell, so a
refined sky kept the model's blocky outline no matter how the control was
set. A marker-based watershed now puts each contour onto a real edge, and
the refinement is a per-layer slider.

Coverage 70.4% -> 70.6% (127/180).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 21:33:16 +02:00
dtourolleandClaude Opus 5 dabf63ed6d Regenerate the matrix over the watershed and sidecar work
`Traceability` has been red on master. The matrix records every tag as
`file.rs:LINE`, so it goes stale two ways at once, and both happened here:

  - the `cargo fmt --all` sweep (710fcbc) moved lines under tags that did
    not otherwise change, which is drift with no change of meaning;
  - the segmentation and mask-storage work added tags and a requirement,
    which is drift that means something.

Regenerated: 332 -> 334 files, 1058 -> 1086 tags, 179 -> 180 requirements
defined, coverage 70.4% -> 70.6% (127/180).

No code changes. The pre-commit hook regenerates this file whenever a
taggable source file is touched, so the way it gets stale is a commit made
with `--no-verify`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 21:32:12 +02:00
dtourolleandClaude Opus 5 f6c9343bcc Ask the pixels where the edge is, not just what belongs
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m53s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 1h18m38s
Build and test / Layer separation (push) Successful in 46s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 3s
Traceability / Requirement traces (push) Failing after 50s
Build and test / Android (aarch64) (push) Failing after 30s
The colour gate decided *what* was in a category and had no way to decide
*where* its edge fell. A colour test has no notion of an edge. So a refined
sky lost its flag and kept the model's twenty-pixel-blocky outline, and no
setting of the control could move that outline onto the horizon.

This adds the second half: a **marker-based watershed**. The mask is eroded
to give two markers, and the flood runs in the ribbon left between them,
meeting along the most expensive line it can find. The cost is a sum of terms
exactly as docs/segmentation.md §2 specifies — the photograph's own edges, and
the colour model's disagreement.

## Why this is not the watershed §15 threw away

That path failed because the merge *ladder* collapsed: 45,808 basins reduced
to one region plus specks. There is no ladder here. Markers prevent
over-segmentation by seeding rather than by merging afterwards, so the one
component that broke is the one component this does not have.

The markers are also better than the textbook's. scikit-image derives them by
thresholding the gradient — guessing where objects are — where these come from
a model that knows what sky is. Marker selection is what normally goes wrong
with this method, and it was already solved.

## The gate still runs, and it runs first

A flood cannot replace the colour gate. It only refines contours that already
exist, and there is no contour around a flag precisely because the model never
noticed one — the flag in the tests sits forty-five pixels from the boundary
against a ribbon of six. The tests caught this; the first version of this
commit had the flood standing in for the gate and the flag stayed.

So the gate goes first and *creates* the contour, and the flood then puts every
contour — the horizon and the new hole alike — onto a real edge.

## Erosion that does not delete flagpoles

Eroding by a cell and a half destroys anything thinner than three cells: a
mast, a bare branch, and equally a strip of sky between two of them. Those
would be left unseeded and the flood would fill them from whichever side
surrounds them, so a flagpole would come back — and come back *confident*.

Erosion therefore stops at the ridge of the distance transform. Whatever would
otherwise vanish keeps a one-pixel seed down its centre, floored at
`min_thickness` so a hot pixel does not qualify. That floor also moves the
signal-versus-noise decision out of colour space, where it was a share of a
fitted distribution nobody can picture, and into image space, where it is a
width in pixels a photographer can see.

## Two modelling errors the outward test found

Both were invisible while the refinement could only subtract, because the gate
was multiplied by weights that were already zero outside the mask. The moment
the boundary could move outward they decided the answer.

**A diagonal covariance is wrong along a gradient.** Sky moves along all three
opponent features together — luminance up, red-green drifting, blue-yellow
down — so treating them as independent charges a colour two deviations along
that gradient three times over. Measured: sky fifteen rows past the sample
scored 11.6 against a threshold of 11.34, so the model refused the very thing
it was refining. The fit now carries a full 3x3 covariance, inverted by
cofactors rather than by a dependency (D13, the NDK).

**Eroded seeds understate the spread, always, in a known direction.** The
sample is drawn from the middle of a category and never from its edge, so for
anything with a gradient the colours nearest the boundary are exactly the ones
left out. The broad mode is therefore fitted wider than its sample by
`SHOULDER`. Same pixel: Mahalanobis 5.9 uncorrected, 1.5 corrected — the
difference between refusing the horizon and reaching it. Only the broad mode
is widened; the tight ones are what discriminate.

## What was given up

Strict subtractivity. It bounded the damage and kept `scene.rs`'s partition
true for free, and it had to go: a mask that may only shrink can sharpen a
horizon inward but never outward, so wherever the coarse contour sat inside the
true edge, the error survived every setting of the control.

The travel bound replaces it. Everything beyond the ribbon is already a marker,
so the flood never reaches it — not "can only remove" but "can only move this
far", and the distance is the model's own uncertainty. That single bound also
retires the connectivity test, the reachability radius and the separate
additive path that an outward-growing rule would have needed. A blue car below
the horizon cannot be gained, not because a rule forbids it, but because the
flood is never there.

`the_colour_gate_only_removes` keeps the older property where it still holds;
`the_flood_cannot_travel_further_than_the_ribbon` holds the new one across the
whole travel of the control.

## Cost

The flood visits only unlabelled pixels, so confining it to the ribbon is not
an optimisation added on top — it is what a seeded flood does. A ribbon of a
few tens of pixels around one contour is a small part of a proxy.

The distance transform is no longer cached, because it has to be measured from
the mask as the gate leaves it and the gate moves with the control. That is one
transform plus one flood per change of the control, against a precompute that
runs the model once.

Verified: fmt clean, clippy --workspace -D warnings clean, 63 dr-segment tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 21:29:12 +02:00
dtourolleandClaude Opus 5 710fcbc1bd Format the face work with the workspace's own rustfmt
Not authored in this session. `cargo fmt --all` reformats every crate, so
running it while working on `dr-segment` picked up five files from the recent
face and library work that had been committed unformatted.

Committed on its own rather than swept into the change that happened to
produce it: the diff is pure whitespace, and mixed into a commit that alters
an algorithm it would be noise in exactly the place someone is trying to read
carefully. `cargo fmt --all -- --check` is a CI gate (tools/ci-local.sh), so
this had to land somewhere regardless.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 21:29:12 +02:00
dtourolle 89c4ff1820 Store what the model found, so a reopened photograph keeps its masks
A subject or category layer was written to the sidecar as identity alone —
which run, which instance, which category — on the reasoning that the pixels
are reproducible by running the same model over the same image. They are, but
only by *running the model*, and nothing runs one except a photographer
pressing "find subjects". So on every path that did not already have a run in
memory the layer resolved to no coverage, `MaskPass::render` logged "has no
distance field; skipping", and the adjustment was silently absent:

  - reopening an edited photograph rendered it without its local adjustments,
    and then saved that state back on the way out;
  - a batch export from the grid could not have them at any point, because
    `render_from_library` opens a session, applies a version and renders, and
    there is no model anywhere on that path. Three hundred files written
    without the edits their photographer made, over a log warning.

Neither failure announced itself. The generated shader still emits the layer's
block and the empty placeholder multiplies it by zero, so the result is a
well-formed frame that is simply missing an edit — `mask_is_stale` already
named the state and called it "not stale, just unrenderable".

The coverage now travels in the file, as one `coverage = w h levels payload`
line at the end of the layer's block.

Two levels, and that is not a compromise. The model hands out a byte per pixel
but `Shaped::build` measures its distance field from `coverage >= 128` and
throws the shoulder away on the first line; everything soft about the rendered
edge comes afterwards from the layer's feather and falloff, which are read off
the distance. So one bit per pixel is not an approximation of what the model
said — it is exactly the part of it that reaches a pixel, and the stored mask
renders the identical frame. Storing all 256 levels would have stored 1.7 MB
of bilinear interpolation to reconstruct a predicate, and would not even have
compressed: a model mask is a bilinear upsample of a coarse grid, so almost no
two adjacent bytes are alike. Measured on a simulated sky and a simulated
figure at 1600x1067, against 1.71 MB raw: 4.0 kB and 6.5 kB at two levels,
46 kB and 76 kB at sixteen, 835 kB and 1.43 MB at all 256. The level count is
still written into the line, so a later build that finds a use for the
shoulder can write sixteen and this one will read them rather than misreading
a stream of lengths as pairs.

The coder is hand-rolled — run-length pairs in a base-64 varint — because
`dr-pipeline` links nothing, which is the property that lets the descriptor
and codegen logic be tested without a device. `flate2` would have been fewer
lines and a dependency in the one crate that has none.

Where it lives matters more than how it is coded. The raster sits on
`MaskLayer` beside the source, not inside `MaskSource::Subject`: the source is
*identity*, which is what makes it diff as a handful of numbers and merge per
field under FR-NC-9, and a raster in there would have given the merge a binary
blob to arbitrate. It takes no part in `MaskLayer`'s equality for the same
reason — a device that has run the model and one that has not hold the same
edit, and counting the difference would raise a conflict over a cache and let
`remote_wins` answer it by discarding the only copy of the pixels.

Encoding happens in `masks_for_storage`, on the save path, rather than in
`ensure_subject_fields` where every coverage already funnels through.
`ensure_subject_fields` runs on a drag — dilating a mask with a compound
morphology rebuilds the field every frame — and encoding a megapixel raster
per frame is the kind of work NFR-P5 exists to keep off a gesture. Saving
happens once, when the photograph stops being the open one, and already costs
a network round trip.

Version skew holds both ways. A file with no `coverage` line reads exactly as
it did before, which is a layer that needs the model run; an unreadable one
costs the pixels and not the layer, because the layer is the edit and the
raster is a cache of it. An old build reading a new file drops the key it does
not understand, which costs a model run and no work. And a payload that will
not compress is refused rather than truncated: a checkerboard would encode to
twice the raster it came from, so past 64 kB nothing is stored and the
behaviour falls back to what it was — half a mask would render as a mask that
is confidently wrong, which is the failure that tells nobody.
2026-08-30 21:28:54 +02:00
dtourolle 6acc98baad Carry the photograph's header into an export made from develop
The same file exported from the library grid kept its camera, its lens,
its capture date and its rights statement. Exported from the develop
button it kept none of them, and `{date}` in a filename template
resolved to nothing at all. Two buttons, one photograph, two different
files -- and the develop one was the version the photographer had just
finished working on.

A session now remembers the header it was opened from, and
`open_session` takes that header rather than the orientation read out of
it, so a photograph cannot be opened for editing without saying which
file it came from. `render_open_frame` clones it onto
`Source::Rendered`; both arms of `export_one` -- the worker's own decode
and the frame handed over already rendered -- turn a header into a
`{date}` and a `SourceMetadata` through the same function, so the two
paths cannot come to different readings of one file. What of it actually
reaches the exported bytes is still decided inside `dr-export` from the
settings, which is what keeps the location-stripping option working here
rather than giving it a second implementation to disagree with.

The alternative was to hang the metadata on `Source::Rendered` alone and
keep it beside the session in the interface. That touches less, but it
makes the header and the pixels two cells to hold in step across the six
places an image is opened, replaced or fails to open, and the failure
mode of getting that pairing wrong is not a missing tag: it is one
photograph exported under another's byline and coordinates, silently.
Kept on the session, the two travel together or not at all.

The header is stored decoded rather than transcribed at open time,
deliberately. `dr-export` argues that source metadata is a parameter and
not a field on `Frame`, because two exports of one frame may legitimately
disclose different amounts; by the same reasoning a session may remember
where its pixels came from without that being a decision about what to
publish, and the allowlist that decides remains the single function in
`export.rs`.

A file with no header is left with none -- an empty `{date}` and nothing
for the encoder to copy -- rather than today's date standing in for a
capture time nobody recorded.
2026-08-30 21:28:31 +02:00
dtourolleandClaude Opus 5 353382c07f Hand the photographer's place between devices
A place recorded on the tablet should be where the desktop opens.

Exchanged through `.darkroom-derived/place.json`, beside the thumbnail
shards and the catalog snapshot. Newest timestamp wins outright: unlike
the catalog this is replaced rather than merged, because two devices
cannot both be where the photographer is and so there is nothing of
theirs inside ours to preserve.

It still refuses to upload over a copy it could not read, for a smaller
version of the reason `sync_catalog` does: a record we have not compared
against may be the newer one, and overwriting it would move the other
device's photographer without ever having seen where they were.

Last in the pass, and its failures are logged rather than reported.
Everything else in that folder is *derived* -- a faster way to learn what
the device could work out for itself -- so losing it costs time. A place
is a fact only the other device knew, and losing it costs a scroll. A
sync that ran out of connectivity should spend what it had on the shards.

The full pass runs after a thumbnail sweep or when Sync is pressed,
neither of which happens on an ordinary launch -- so a handover would
arrive one launch late, which is one too many for a feature whose whole
claim is picking up where you stopped. `spawn_place_fetch` is the small
half: one GET of a few hundred bytes, started beside the scan.

And it can still be refused. A handover is welcome on the way in and
unwelcome once the photographer has started: a grid that jumped
elsewhere mid-scroll because a round trip finally landed would have lost
their place to the feature meant to keep it. Any scroll, scrub, scope
change, filter or opened photograph closes the latch, and a record
arriving after that is written to disk and takes effect next launch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 20:40:22 +02:00
dtourolleandClaude Opus 5 d44bffa4a8 Remember where the photographer was
Opening the application was always a fresh arrival at the beginning of
the library, whatever you had been doing when you closed it.

What is written down is the view, the scope, the rating filter and the
photograph on screen -- the open one in develop, the first visible one in
the grid. Not just a scroll position: a position without the filter that
produced it names a row of a list that no longer exists. Restoring them
has an order for the same reason -- scope, then filter, then position,
then the view -- because each step changes what an ordinal *means*.

Addressed by remote path and collection UUID, never by an ordinal or a
row id. `images.id` and `collections.id` are local to one catalog, and a
grid ordinal is local to one ordering; a record naming either would land
somewhere arbitrary on a second device and after any filter change on
this one. Where the ordinal is needed, `library::ordinal_of_path`
computes it through the grid's own `ORDER BY`, taken verbatim by a window
function rather than spelled a second time as an inequality -- which is
the mistake `grid_order_for` already warns about, and which a manually
ordered collection would make unreadable.

Every failure degrades rather than reports. A collection this device has
not merged leaves the scope at the whole library; a photograph that has
since been deleted falls back to when it was taken, which puts the grid
in the right week; a torn file yields no place and the library opens at
the top. Reopening develop is the one thing that requires an exact match,
because a canvas on a path that no longer resolves is a filename over an
empty frame.

The record lives in `dr-types` beside `Settings` and the store lives here
beside `SettingsStore`, for the reason `dr-types`' manifest gives: a JSON
serialiser in `core/` would be paid for by every crate there. Two files
and two lifetimes, though -- resetting preferences must not forget where
you were.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 20:39:50 +02:00
dtourolleandClaude Opus 5 8bf5e13faf Centre the photo roll on the frame it opens with
The roll brought the open photograph into view by the shortest move,
which is right for stepping along it and wrong for the first look: a
frame near either end of the loaded window arrived hard against an edge,
with nothing on that side to give it any context.

It now centres on the first settle of a develop session and steps
minimally after that. A one-shot request that the strip itself clears --
the only thing that knows the request has been honoured is the code
honouring it -- rather than something recomputed on creation, because the
strip is created far more often than a session begins: leaving develop
for Settings and coming back rebuilds it, and re-centring then would undo
a roll the user had scrolled by hand.

Raised on the two ways into develop from the grid, and not on a pick
along the roll, which is a step within a session rather than the start of
one.

Centring is clamped to the ends: the third photograph of a window cannot
be centred without scrolling empty space in beside it, and a strip that
begins with a gap reads as broken rather than as centred.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 20:35:55 +02:00
dtourolleandClaude Opus 5 0ff1e01ec3 Come back from develop on the photograph you were editing
Leaving develop returned to where the *grid* was, which after a walk
along the photo roll can be a thousand rows from the frame you had just
finished. So the one photograph you were certainly interested in was the
one the grid came back without.

Two positions, and the rule is not to pick one of them. The grid seeks to
the remembered position, then reveals the keyboard cursor -- which is now
put on the open photograph, and which moves the viewport as little as
will bring its row into view. A frame inside the remembered screenful
moves nothing at all; one outside it scrolls exactly far enough. One
rule, both behaviours.

The cursor rather than the selection, deliberately: `place_cursor` also
rewrites the selection, and a set of forty photographs assembled in the
grid must survive having one of them opened.

`reveal()` now also runs on the grid's `init`, since `cursor-row` is
initialised rather than changed when the subtree is rebuilt and no
handler would otherwise fire. Both it and the roll's centring defer while
the element has no height yet -- `init` runs before layout, where a
height of zero makes every row look off screen -- and a latch brings the
first real height back to the cursor without letting every later resize
haul the viewport around.

The capture-time marker follows the same move, for the same reason:
`load_window` rebuilds the axis only when the scope, the filter or the
total has changed, and none of them has. It is the same library seen from
a different row.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 20:35:23 +02:00
dtourolleandClaude Opus 5 53dc2d171e Say where the grid is on the capture-time axis from the first frame
The sidebar's marker rested greyed at mid-track until the first scroll or
scrub. The reasoning was that anchoring it would imply a choice the user
had not made -- but that reads the marker as reporting an intention, and
it does not. The sidebar's whole claim is to say *when* you are, and that
is known from the first frame: the grid is at the top of the library, or
wherever it was last left.

So a launch opened with the marker halfway down an axis whose visible
photographs were all from the wrong end of it. Dimmed rather than absent,
which made it look like a reading rather than the absence of one.

Seeded in `refresh_timeline` -- the one place that decides what the
marker says, and the one that runs on every route which builds the axis
-- and only when nothing has claimed it, so a scroll or a scrub still
speaks for itself.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 20:34:11 +02:00
dtourolleandClaude Opus 5 eed27eb36d Keep the grid's place when another screen covers it
Opening Settings, Import or People and coming back landed at the top of
the library however deep in it you had been.

The grid is gated on an `if` in the markup, so every route away from it
destroys the subtree and rebuilds it. A Flickable being destroyed passes
its viewport through zero on the way out, and that reaches
`on_library_scrolled` looking exactly like the user having flung the grid
to the top. The handler already guarded against it -- but on
`show-library`, which means "the library rather than develop" and stays
true while any of those four screens replaces the window. So the guard
covered the develop route and none of the other three: `resume_at` was
overwritten with 0 on the way out, and the position was gone before
anything could restore it.

The condition the `if` is actually spelled with is now computed once, in
`app.slint`, and Rust reads that. The two cannot drift apart again
because there is only one of them.

That fixes the overwrite. The second half is that nothing replayed the
position on the way back in: `on_back_to_library` does it by hand, and
Settings, Import, People and the launch screen do not go through it.
Rather than teaching three more modules to call it, `scroll-to` is now
kept current on every scroll. It is read by `seek()`, which runs on a
token change and on `init`, so writing it without bumping the token
cannot move the grid on screen -- and is exactly what the next grid reads
when it is built.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 20:33:34 +02:00
dtourolleandClaude Opus 5 4dc954f01a Take the photo roll's grab band off the buttons that end a mode
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m53s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 58s
Build and test / Layer separation (push) Successful in 45s
🐳 Android image / Build and push (push) Successful in 4s
Build and test / android-image (push) Successful in 4s
Traceability / Requirement traces (push) Failing after 1m0s
Build and test / Android (aarch64) (push) Failing after 31s
"Done Cropping", "Done Repairing", "Done Masking" and "Fit" float over
the foot of the canvas. So does the photo roll's swipe handler, and a
gesture handler is not a layout box — it is an input surface. A press
inside one is delayed, then offered to that handler's own children and
to nothing else: `input_event_filter_before_children` returns
`DelayForwarding`, which aborts the hit-test traversal outright, and the
replay afterwards visits only the handler's subtree. Everything behind
it is never asked, hover included.

The band was `strip-height + reach` — 136px along the bottom — whether
the roll was out or away. So the button that ends a mode was drawn, was
lit, and did nothing for as long as a library was open, which is the
whole time anybody is developing from one. The tool rail kept working
because it is a sibling of the canvas rather than behind the roll, which
is exactly why this looked like two dead buttons rather than a dead
region.

The band now goes where the roll goes. The handler carries the strip
instead of standing still while the strip animates inside it: closed,
only `reach` is on screen and the rest hangs below the window where
nothing can press it; open, it still covers the thumbnails, which is
what lets a swipe down anywhere across them put the roll away. The
180ms travel moved from the strip onto the handler, so the drawn
positions in both states are what they were.

The controls are then positioned against that band rather than against
the bottom of the canvas, and ride up with the strip when it comes out.
Reordering them in front of the roll would have been the other fix, and
it is the wrong one — the band would become the thing that cannot be
reached, and a gesture nobody can start is worse than a button with a
second way out.

`roll-strip` and `roll-reach` are tokens now, because two files have to
agree on where that band is for either of them to keep out of it.

The bottom of the photograph comes back with it: the crop's lower
handles and a repair placed near the bottom edge were inside the same
136px and had the same fault.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:42:17 +02:00
dtourolleandClaude Opus 5 3d248cfb79 Export the photograph, not the canvas
Zooming the develop view changed the exported file. `Framing::view` is
kept out of the sidecar, out of `is_active` and out of `output_size`
precisely so that it cannot — but those exclusions keep it out of the
*edit*, and an export is a *render*. `visible_rect` deliberately folds
the view into the single rect the fused shader's prologue samples, so
`render_for_export` inherited it: at 4:1 it wrote the middle of the
frame, magnified to fill the file at the full output size, with the
detail kernels scaled four times over because `render_scale` folds the
view in as well. `render_thumbnail` did the same to the grid.

`render_uncropped` already suspends the view for this exact reason, so
the fix is its pattern: one `render_the_file` that both file-producing
paths go through, composing inside the suspension since the view
reaches the shader as a uniform baked at composition time. Restored
whatever happens — leaving the graph un-zoomed after a failed export
would throw away where the photographer was looking.

Nothing caught it because the guard checked the wrong things.
`zooming_does_not_change_the_exported_image` asserted the output size
and the crop; both held perfectly throughout. Renamed to
`zooming_does_not_change_the_size_or_the_crop`, which is what it
tests, and the pixels are now guarded where pixels exist. The new test
uses a ramp rather than quadrants deliberately: a four-quadrant frame
is self-similar under a centred zoom, and the first version of this
test passed against the bug because of it.

Traces FR-EXP-9, which asks for the full-quality pipeline "regardless
of what the display was showing".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:41:30 +02:00
dtourolleandClaude Opus 5 a1e361e35a Measure the native path against the one it replaces
Everything argued for this change so far was read out of a catalog
after the fact: crop_px across 18,671 faces the old code had already
stored. That is evidence about what the previous implementation did. It
is not evidence that the new one does better, and the difference matters
because a landmark left in the detector's coordinates, or a box filter
with an off-by-one in its source span, would both produce faces that
look entirely plausible until somebody counted the pixels behind them.

So examples/face_native.rs renders one file and indexes it twice, native
and from a 1024 proxy, changing nothing else. Fourteen originals from
the reference library, 5472x3648 CR2 and DNG:

    native      9 faces, mean crop 287px
    1024 proxy  5 faces, mean crop  75px

Crops 3.8x larger, and across the line that decides whether the crop is
photographed or interpolated: 75px is below ALIGNED_EDGE, so the proxy
path was upsampling into the embedder on average where this one
downsamples into it. Fourteen images and nine faces is enough to show a
direction and to catch a wrong scaling; §7b says so rather than quoting
the ratio as a library-wide figure.

It also corrects something §7b asserted two commits ago. I wrote that
detector input resolution cannot affect recall, because §4.1 letterboxes
everything to 640. Native found nine faces to the proxy's five,
including four on files where the proxy found none, so it plainly can.
The two paths differ in their resampling as well as their size, and this
experiment does not separate those, so §7b now records the result as
evidence for the double-resampling hypothesis rather than as its proof.
M4 still owns settling it.

The audit-summary test went stale when the ready/to-fetch split was
collapsed and is updated to assert the single number, including that the
old wording is gone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:41:30 +02:00
dtourolleandClaude Opus 5 4af3b93dfa Index faces from the native render, not from a preview of it
Implements the FR-CULL-8 written two commits ago. The sweep fetched the
JPEG preview embedded in each RAW and used that one buffer for both
detection and the crop; it now fetches the original, renders it through
the same path export uses, reduces that for the detector, and warps the
crop back out of the native frame.

Three pieces, and each exists for a reason worth stating.

dr_face::Pixels lets the warp sample 8-bit RGBA directly. A 24 MP native
frame is 96 MB as RGBA and 288 MB converted to the f32 RGB align.rs was
written against, and the warp reads about forty thousand pixels out of
it. Converting the whole frame to sample 0.2% of it is NFR-RES-2's
budget spent on a copy, per image, for a whole library. The variant
costs one branch per sample and a test asserts both layouts produce
identical crops.

The detector gets a box-filtered reduction to 1600px, not the native
frame and not a point-sampled one. Averaging rather than sampling
because the detector's job is finding small faces and decimation is
precisely the operation that removes them: at 4x, fifteen of every
sixteen pixels are discarded and a 40px face survives or not depending
on where it falls relative to the sample grid. 1600 rather than 640
leaves the letterbox a mild 2.5x rather than a 9x, and bounds the f32
buffer at 20 MB.

Landmarks come back in the reduction's coordinates and are scaled to
native in one place before any crop pixel is read. This is the failure
mode that would not announce itself -- unscaled landmarks put every crop
near the top-left corner, which yields faces of something else, cleanly
embedded and confidently clustered.

The sweep fetches SWEEP_LANES-wide and renders sequentially. Not a
placeholder for a parallel version: there is one GPU, so concurrent
renders queue on it regardless, and each materialises a native frame.
Overlapping them would multiply the one allocation that threatens the
memory budget while buying parallelism that does not exist. The chunk
drops from 96 to 6 for the same reason -- 96 held 8 MB previews, this
holds whole RAWs.

The stored edit is deliberately not applied, which is where this departs
from export::render_from_library. Face geometry is normalised to the
frame, so indexing a cropped render would record boxes against a frame
that changes whenever the user changes their mind, and every stored box
would quietly become wrong. Orientation is applied: that is a fact about
the file rather than an edit.

examples/face_native.rs renders one file and indexes it both ways, so
the claim behind all of this can be checked against photographs rather
than re-read out of the catalog it came from.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:41:30 +02:00
dtourolleandClaude Opus 5 9ddc1273c0 Make a sweep that fails everything say so
A run over 169 images failed all 169, in fourteen seconds, and reported
"0 face(s) in 0 image(s)" -- the same sentence a run that indexed
nothing because there was nothing to index produces. Three separate
places dropped the information on its way to the screen.

The progress count only moved on success. FaceSweepMessage had no
failure variant at all, so a pass where every image failed sat at 0/169
from the first tick to the last: the receiver was told the total, told
nothing, and told the pass had ended. That is indistinguishable from a
hung job, and it is what it was taken for.

Finished already carried a failed count and identity_ui matched it with
`Finished { .. }`, throwing the number away and printing the tidy
success line regardless.

And the reason each image failed was logged at debug, which is off, so
169 consecutive failures left no trace of why anywhere.

Failed { images } now carries the count back per lane batch, the
progress counter advances on it, and both the running status line and
the finishing activity row say how many could not be read. A batch
rather than one message per image because failures come back lane-sized
and the useful number is how many.

Also renames the store sweep's guard to MIN_CROP_EDGE with the rest of
that constant's move, since the two touch the same lines.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:41:20 +02:00
dtourolleandClaude Opus 5 e7b526c550 Specify face indexing at native resolution, and say what the proxy cost
FR-CULL-8 said detection runs against the thumbnail or proxy tier and
never a full decode, and faces.md §5 said the aligned crop is sampled
from that same proxy. Both are wrong in the same place: they treat
detection and cropping as one resolution problem when they are two, with
opposite answers.

Detection does not care. §4.1 fixes the graph's input at 640x640 and
letterboxes whatever arrives, so a face filling 2% of the frame reaches
the model at 12px whether the buffer handed over is 1024px or 6000px.
Every pixel above the detector's own input is discarded before inference.

The crop cares about nothing else. §5's warp produces the fixed 112x112
ArcFace sees, so source resolution converts directly into whether those
112 pixels were photographed or interpolated. Reading crop_px across the
18,671 faces the proxy-tier implementation stored: 47.3% were upsampled
to reach the embedder, 314 of them by more than 2x, the smallest from 34
source pixels. An upsampled crop does not fail loudly -- it yields a
confident embedding of detail that was never there, and the damage
appears three stages later as clusters that will not separate.

So FR-CULL-8 now specifies four stages with the resolutions named
separately: render native through FR-EXP-9's pipeline, downscale for the
detector, map boxes and landmarks back to native, crop and align from
the native render. The affordability the old rule bought is met instead
by when the pass runs -- background, preempted, resumable -- and the
requirement says plainly what it now costs on a remote library: the
original rather than FR-NC-3's byte range, 412 GB across the reference
library's 19,107 images, so a whole-library pass is a transfer under
FR-NC-6 rather than something that may start on its own.

MIN_CROP_EDGE replaces the MIN_DETECT_EDGE this branch briefly had. Same
number, guarding the quantity that turned out to matter.

faces.md §7b records both measurements, and marks the second as
unexplained rather than dressing it as a finding. Grouped by the buffer
detection ran against, faces per image was 0.078 at 1024 or below and
1.82 at 2048 or better, controlled for file type and size. That gap is
real and reproducible and I cannot account for it, because the letterbox
above says detector input should not matter. M4 is where it gets
settled. The crop measurement does not depend on it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:41:20 +02:00
dtourolleandClaude Opus 5 144d2e4e84 Revert the detection floor: it guards the wrong resolution
Reverts 53f7cdf and e92d22d. The floor those added sat on
Detector::detect, refusing any buffer under 1025px on the reasoning that
a small buffer finds no faces. That reasoning does not survive §4.1:
the detector letterboxes every input to 640x640, so a face occupying 2%
of the frame presents at 12px to the model whether it is handed a 1024px
buffer or a 6000px one. Detector input is precisely the quantity that
does not matter.

Worse than merely useless, it blocks the design FR-CULL-8 now specifies,
where the detector is deliberately fed a downscale and the crop is taken
from the native render. A guard on detect() rejects exactly that call.

What the measurement actually supports is a floor on the *crop* source,
which is where resolution converts into embedding quality, and which
faces.crop_px already records: 47% of the reference library's faces were
upsampled to reach 112x112. That floor is a separate change against the
native-resolution path and does not belong on the detector.

The 23x faces-per-image gap by source_edge that motivated the original
commit is kept in faces.md §7b, restated as the unexplained observation
it is rather than the causal claim it was written as. V12 stands: those
runs cropped at 1024 whatever detection did, and that is reason enough
to look at them again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:41:12 +02:00
dtourolleandClaude Opus 5 41c655c176 Stop the local-store pass pretending it can still index
spawn_store_face_sweep detects on what the thumbnail store holds, and
the store's largest tier is FACE_TIER -- 1024, which the floor now
refuses. Left alone it would list every outstanding image, decode every
proxy it had, and report every single one as failed: twenty thousand
refusals all saying the same thing, with the real explanation buried at
debug level.

There is no repair available inside this pass. It has no larger tier to
read; the pixels the detector needs have to come off the server, which
is spawn_face_sweep's job and always was. So the honest behaviour is to
check the tier against the floor once, say plainly which pass to use
instead, and stop. It is only reached from examples/face_index.rs, so
this costs the tool its --run mode and no shipped behaviour.

FACE_TIER keeps its value and loses its meaning. It is now the tier a
stored crop is *cut from*, which 1024 is entirely adequate for -- the
face has already been located and the crop only has to be looked at --
and no longer the tier faces are *found* on, which is the thing that was
returning 0.078 faces per image.

IndexAudit's ready/awaiting_proxy split goes the same way. It existed
because one of the two passes could only do images that already had a
proxy; now that detection refuses that proxy's size, both halves cost
the same fetch, and a status line reading "169 ready to index" implies a
distinction that no longer decides anything. One number.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:41:04 +02:00
dtourolleandClaude Opus 5 d0ebc9f571 Count the same images in the progress figure that the sweeps index
The Identity screen said 4,593 images were left to index and stayed
there for hours across repeated runs, which is what a stuck job looks
like. It was not stuck. 4,424 of those 4,593 are shadowed -- the JPEG
half of a RAW+JPEG pair -- and no sweep will ever index one, because
every work list is built on VISIBLE, which excludes them. They are not
separate photographs and the grid does not show them either.

But faces::coverage counted them: its denominator was "images WHERE
trashed_at IS NULL", with no shadowed_by clause. So the outstanding
figure had a floor of 4,424 that no amount of work could bring down, and
Coverage::is_complete could never once return true no matter how
completely the library had been indexed. A progress number that cannot
reach its own target is worse than no progress number.

The fix is to count the population the sweeps actually draw from, in all
three places that were describing it differently: coverage's denominator
and its indexed join, and audit's split of the outstanding set, which
had the same gap and fed the same status line.

On the reference library the denominator goes from 23,531 to 19,107 and
outstanding from 4,593 to 169 -- the second of which is a number the
user can watch go down, and which turns out to be a real and separate
fetch failure worth chasing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:41:04 +02:00
dtourolleandClaude Opus 5 6783d0c723 Settle the images whose best preview is under the floor
The floor introduces a state the sweep had no arm for. An image whose
largest embedded preview is genuinely smaller than 1025px now comes back
from index_preview as an error, lands in the generic Err arm, and is
counted as failed -- which means no face_index row, which means it is
still outstanding, which means the next sweep fetches exactly the same
bytes and refuses them again. For ever, on every run, at one range
request each. The previous behaviour was wrong but at least terminated;
this would not.

So ProxyTooSmall gets its own arm, and it records a marker at the true
edge rather than nothing. That is the difference between "we looked and
found nothing" -- which would be a lie, since nothing was looked at --
and "this was examined at 900px, which is the best this file has". The
first is unrecoverable; the second is a fact source_edge was added to
carry, and a later floor or a bigger proxy can select on it deliberately
the way V12 just did.

Counted separately from failures all the way up, because they are not
failures and reading them as such would misdescribe a library of small
scans as a broken network. The summary line says how many and why.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:41:04 +02:00
dtourolleandClaude Opus 5 0a6509de93 Forget the face runs made on proxies too small to see a face
The floor added in the previous commit stops this happening again; it
does nothing about the 1,824 images in the reference library that
already carry a face_index row written against a proxy of 1024 or less.
Those rows are why the damage is permanent rather than merely past. The
work list is "images with no row for this model", so an image examined
against a 1024px proxy -- 0.078 faces per image, nine in ten finding
nothing -- is indistinguishable from one examined properly, and no
later pass will ever offer it to the detector again.

V12 deletes exactly those markers, and nothing else. The faces those
runs did find stay in place and keep drawing the People screen until a
better pass replaces them, and record_detections re-attaches the user's
confirmed names across that replacement by box overlap, so a library
somebody has spent an evening naming does not lose that evening. The
cost is a re-fetch of the affected images.

Deleting the marker rather than teaching the work-list query to select
on source_edge, which was the other option and is worse. A standing
`source_edge < floor` predicate never lets go: an image whose largest
embedded preview is genuinely smaller than the floor would be re-fetched
on every sweep for ever, because the next pass cannot do any better than
the last one did. A one-off deletion gives each affected image exactly
one more attempt through the good path and then lets the ordinary
"has a row" rule settle it.

The threshold is written out in the SQL instead of referring to
dr_face::MIN_DETECT_EDGE. A migration has to keep meaning what it meant
when it ran; binding it to a constant someone may raise later would
quietly change what an old catalog gets migrated to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:40:26 +02:00
dtourolleandClaude Opus 5 1d800d56b0 Refuse to detect faces on a proxy too small to find one
Detection was run on whatever proxy the caller happened to have. A
small one does not fail -- the image is letterboxed into the detector's
640px input at any size -- so it comes back with almost nothing, and the
caller then writes a face_index row saying the photograph was examined.
That row is the damage. Nothing distinguishes it from "examined
properly, no faces in this one", so the image is never looked at again.

The measurement, on the reference library of 23,531 images. Runs against
a 1024-edge proxy: 0.078 faces per image, 90% of them finding nothing at
all. Runs against 2048 or better: 1.82. To rule out the obvious
objection that small proxies just come from small photographs, the same
comparison restricted to DNGs -- 1,592 of them averaging 21 MB against
7,724 averaging 23 MB, so the same kind of file in the same library --
gives 0.078 against 1.82 again. Twenty-three fold, on identical source
material, identical weights, identical options.

So the floor goes in the detector rather than in either sweep, because
both of them, the example tool and any future job handler are equally
entitled to get this wrong, and there is one place that sees every
attempt.

It is 1025, not 1024, and the odd-looking number is the point: 1024 is
exactly ThumbSize::Large, the tier proxies are stored at and the tier
one of the two sweeps was detecting on. A floor that admitted 1024
would admit precisely the population this exists to exclude. Written as
a minimum rather than a maximum so the test at each call site is
`edge < MIN_DETECT_EDGE` with no boundary left to get wrong.

ProxyTooSmall is its own error variant rather than an empty result
because the caller has to tell it apart from a failure: nothing is
wrong with the image or the model, and the answer is to go and find
better pixels, not to retry these ones.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 19:40:12 +02:00
dtourolleandClaude Opus 5 7981718d83 Time the phases of a launch, because the tablet has no profiler
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m8s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 1h19m46s
Build and test / Layer separation (push) Successful in 47s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Traceability / Requirement traces (push) Failing after 46s
Build and test / Android (aarch64) (push) Successful in 1h1m21s
Two things are now off the launch path and the rest of it is unmeasured.
There is no way to attach a profiler to an Android launch, and the
window that matters — from `android_main` to the first `poll_events` —
is over before anything on the device can be asked a question about it.
So a line in the log file is the only measurement anybody gets.

Three of them: the GPU open, the window build, and the total to the
event loop. The last is the one that matters, because it is the figure
the input dispatcher is counting against — anything approaching five
seconds there is the next ANR whatever the phases above it say.

The GPU open is timed rather than moved. It is a Vulkan instance, an
adapter enumeration and a device request, and on the desktop it cannot
be deferred at all: it selects the Slint backend, and creating a window
selects one for us. On Android it could be, because nothing shares that
device with the compositor (TD-1) — but "could be deferred" is not
"costs enough to be worth deferring", and there is no number yet that
says which. This is the line that will produce one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:41:34 +02:00
dtourolleandClaude Opus 5 d48e9f6033 Open the catalog on a worker, so a launch is not a page-by-page read
The second thing standing between `android_main` and the first
`poll_events`, and the one that grows with the library rather than with
the APK.

`library_ui::open` called `show_catalog_now`, which called
`Catalog::open_verified`. That runs `PRAGMA quick_check`, which reads
every page of the database, and then `Catalog::open`, which takes a full
SQLite backup of the file before a migration and rewrites its structure
afterwards. On a 50,000-image library that is tens of megabytes of I/O
on a tablet's flash, and it happened before the window had painted
anything — so on Android it was counted against the five seconds the
input dispatcher allows, and on the desktop it was a launch that sat on
a blank window.

`dr_catalog::recovery`'s own module documentation says the check is
affordable "at startup, where a failure has a user in front of it who
can answer a question". That was the intent and it was not true: there
was no interface yet in which to ask. Now there is, because the open
happens on a worker and the answer arrives on a channel drained by a
timer — the same shape the scan, the thumbnails and the login already
use.

The gate the synchronous call provided is kept, and is the reason the
scan moved with it. `Catalog::open` succeeds on a damaged file whose
header survived, so a scan running beside an unanswered recovery
question writes ETags and image rows into damaged pages and turns a
catalog that had a backup into one where the backup is the only copy
left. So the scan now starts from the drain, on the two answers that
permit it, and not at all on `Corrupt`. `library-scanning` stays true
throughout, which hides the Rescan button and stops the gate being
merely advisory.

What the user sees while it runs is a third empty state. The grid
already refused to conflate "still scanning" with "scanned, found
nothing"; "opening the library" is a third answer and it gets its own
sentence, because a grid saying "Scanning…" while nothing is on the
network is the same kind of lie the other two were separated to avoid.

`show_catalog_now` stays, unchanged and blocking, for `recovery_ui`.
That call site has the event loop running, has just replaced the file
under a `forget_catalog`, and has `recovery-busy` on screen — the same
reasoning `recovery_ui::answer` already gives for doing its file copy in
place. The part both paths share is now `adopt_catalog`.

One consequence worth naming: the cache-usage figure on the settings
page was read at startup from a catalog that is no longer open by then.
It moves to the page's `on_open` closure, beside the face coverage,
which is read there for exactly the same reason — it is only ever looked
at while that page is on screen.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:41:34 +02:00
dtourolleandClaude Opus 5 1c5849ebe8 Unpack the bundled models on a worker, not on the way to the first frame
Launching v0.10.0 on the tablet produces an ANR: "Waited 5000ms for
MotionEvent", 5,827 ms on one input sequence, over a window that has
never painted. The app recovers and sits at 0% afterwards, so it is a
startup cost rather than a hang.

The structural fact behind it is that `android_main` runs with the
activity's input channel unserviced. Nothing drains it until Slint
reaches `poll_events`, and Slint does not reach `poll_events` until
`dr_ui::run` calls `window.run()` on its last line. Every millisecond
before that is a millisecond the input dispatcher waits on, so five
thousand of them is an ANR whatever the work happens to be.

The largest single piece of that work was here. `install_bundled_models`
copies 41 MB on the first launch after an install — 24.9 MB of scene
model, 13.6 MB of embedder, 2.5 MB of detector — each read whole out of
the APK into a `Vec` and written to `/data`, in a loop, on that thread.
v0.10.0 is the release that added the scene model, which is 60% of that
total, and it is the release the ANR appeared in. The 8,010 minor faults
in the report are about what 41 MB of freshly touched pages costs.

So it moves to a detached thread and the function returns as soon as the
thread is running. Nothing on the launch path wanted the result: the
only two things that read these files are the People screen and the
scene tab, both of which are reached by hand, minutes later, from
workers of their own.

What that costs is a window in which a model looks absent.
`library::face_models` and `library::scene_model` decide availability on
`is_file()`, so during the copy both report their feature unavailable —
which is the same answer they give a build shipping no weights at all,
the ordinary case both were written around. Briefly pessimistic rather
than wrong, and the temporary-name-then-rename that was already there is
what keeps it from being worse than that: a lookup never sees a
half-written file, only an absent one. Both call sites now say so.

A completion line reports the bytes copied and the milliseconds taken,
including when it is zero, so the second launch after an install can be
told from the first in a log rather than by inference.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:41:34 +02:00
dtourolleandClaude Opus 5 4efab496c2 Put the refinement on a slider, per layer
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m27s
Benchmarks / Frame budget (on demand) (push) Skipped
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Desktop (Linux) (push) Failing after 1h14m30s
Build and test / Layer separation (push) Successful in 43s
Traceability / Requirement traces (push) Failing after 46s
Build and test / Android (aarch64) (push) Successful in 1h8m56s
The evidence was being gathered and spent immediately at one strictness
nobody could see or change. This makes it a control.

`MaskLayer::refine` is shaping, like the feather, and it lives on the layer
for the layer's reason: two layers may sit on the same category and want
different amounts of it, and the model ran once for both.

## What the segmentation now stores

`CategorySummary` keeps the **coarse** mask and the `Refinement` beside it,
rather than a refined mask. That is what gives the control an off position
that is bit-for-bit the model's own weighting, and what stops a strictness
change from needing the model again.

`category_mask_at` borrows at zero and wherever no refinement could be
fitted, so a layer nobody has touched costs nothing over the old path.

The fit moved to the far side of the orientation permutation. The verdict is
a per-pixel field over the same grid as the mask it gates, so fitting it
upright would mean permuting a proxy-sized buffer afterwards to match — a
second rotation, and a second chance to get one wrong. One logit cell is the
same number of pixels either way: the letterbox scales by the longer edge and
a permutation does not change which edge that is.

The two frames became a `Frames` struct rather than six parameters. This
function reads the picture twice for opposite purposes — the model needs it
upright or it recognises far less, the refinement needs the sensor's grid —
and a transposed pair produces a plausible mask over slightly the wrong
pixels, which is the failure this module is most prone to.

## It rebuilds the field, and the signature says so

`refine` is mixed into `subject_signature`. Unlike a feather, which is read
off a field that is already correct, this changes which pixels are in the
mask at all — so it changes the coverage the field is measured from. Omitting
it is the bug where the slider moves and nothing happens until some unrelated
control invalidates the cache.

That puts it in the same cost class as a close or an open, which is why the
row takes `SliderRow::changed` — already once-per-gesture, since that row
takes `SliderTrack`'s `committed` internally — rather than a live stream.

The slider is offered only where there is something to move: a category
source, *and* a refinement the frame actually gave enough to fit. A control
that moves and does nothing is worse than an absent one.

## Two defaults that are deliberately different

A layer added from the panel starts at 4.0, because a category's edges are
twenty proxy pixels wide before anything is done to them and a photographer
adding a sky mask wants the sky rather than the sky plus every chimney in it.

A layer read from a sidecar with no `refine` key starts at **zero**. A file
written before this control existed has to render as it did then, and a
default of 4 on absence would quietly re-grade every stored category mask in
the catalogue. `a_categorys_refine_strictness_survives_and_defaults_off`
holds both halves, and `an_out_of_range_refine_is_clamped` holds the file to
the scale — past the top of it every colour fails and the mask deletes
itself, which reads as lost work rather than as a bad file.

`MAX_REFINE` is dr-pipeline's own constant mirroring
`dr_segment::STRICTNESS_MAX`, following `Falloff` and `Morphology`: this
crate holds the description of an edit and must not depend on the crate that
runs a model. dr-ui is where the two meet, and the only place that converts.

Verified: fmt clean, clippy --workspace -D warnings clean, 488 dr-pipeline
and 60 dr-segment tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:40 +02:00
dtourolleandClaude Opus 5 f87bf6ebc0 Charge a colour mode for its rarity, and keep the verdict
The refinement worked and could not be controlled. Pruning modes below a
share threshold made the flag's removal a *discrete* event: below the line
its Mahalanobis distance was enormous and nothing rescued it, above the line
it sat at zero and nothing removed it. A control over that would appear dead
through most of its travel and then start eating sky.

So the prune is gone. A mode is charged `−ln(share × k)` nats, floored at
zero, and that cost enters both tests — doubled in the chi-square, which is a
squared distance, and directly in the log density. Rarity becomes a distance
rather than a threshold, and the things a photographer wants to remove
separate along it.

Measured on the synthetic frame the tests build: a flag holding 1.6% of the
sky is more than half gone by **2.95 nats** and a cloud bank holding a third
of it survives to **5.75**. The whole interval between them is somewhere a
control can sit. `the_flag_goes_before_the_cloud_does` pins the ordering,
which is the property that makes one slider worth offering at all.

Measured against an even split rather than against one, so raising
`clusters` describes a category more finely without making every colour in it
look rarer. Floored at zero so a dominant mode earns no *discount* — a
bonus there would let the commonest colour outvote a bad chi-square, which is
the one direction this must not bend.

`Refinement` holds the per-pixel verdict, quantised to a byte over ±16 nats —
an eighth of a nat per step, far finer than the narrowest transition the gate
can be asked for, and the same size as the coverage buffer it sits beside.
`apply` is then a smoothstep, and the model is never consulted again.

That is `distance.rs`'s arrangement deliberately: there a signed distance
field is computed once and feather, grow and shrink become arithmetic on it,
"which is what makes those live controls rather than ones that stall on every
drag". Same shape, different field.

The blur moved with it, from the gate to the verdict. Smoothing the evidence
rather than the decision means it is paid for once in `compute` instead of on
every frame of a drag, and it is the better thing to smooth in any case.

`apply` at `STRICTNESS_OFF` returns the weights untouched without reading the
verdict at all. A control whose off position is *very nearly* the unrefined
mask cannot answer "is this helping"; one whose off position is the unrefined
mask can. `strictness_zero_changes_nothing` holds it to that, and
`strictness_is_monotonic` holds the rest of the travel to only ever removing
more — a slider that gave weight back partway up would be one whose direction
nobody could predict.

The synthetic sky is smooth enough to sit on `VARIANCE_FLOOR`, where a real
one has noise and therefore a real spread, which moves every crossing down
together. The ordering survives that; the placement is a calibration. Which
is the honest argument for a control rather than a constant, and why the
default sits at half scale instead of at the flag's measured crossing.

The example sweeps the whole range and writes a frame per nat, because the
question a photographer asks of a slider is where to put it, and that needs
the travel rather than a point on it.

Verified: fmt clean, clippy -D warnings clean, 60 dr-segment tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:40 +02:00
dtourolleandClaude Opus 5 4f4abd335f Cut a scene category back to the pixels that agree with it
A flag in the sky came out weighted as sky, and no feather setting fixed it.

The scene model's logits are `[1, 150, 80, 80]`, so one cell is eight input
pixels; at the 1600px proxy the letterbox scale is 0.4 and **one cell is 20
proxy pixels**, which `rasterise`'s bilinear then spreads across one more
either side. A flag is a handful of cells whose softmax is dominated by the
sky around it. The information was never in the grid, so nothing downstream
of the grid can recover it.

Tiling is the answer for an instance and is not available here: a category
has no bounding box to tile over — sky is wherever the sky is. But the
photograph is at full proxy resolution even though the weights are not, and
it knows exactly where the flag is. So the model says *what*, and the
pixels say *which of them*, which is the division of labour arm C already
draws between the instance model and the watershed.

## Seeds, and why the erosion radius is not a guess

Threshold the weights high, take `signed_distance`, and keep what is more
than 1.5 cells inside. One cell *is* the model's resolution and the bilinear
spreads it across one more, so the band either side of the boundary is smear
rather than evidence. Deriving the radius from `Scene::cell_pixels` rather
than picking a pixel count means it stays right if the proxy edge or the
export changes.

The mirror of that set is a confident *exterior*, free from the same field.

## Dropping small modes is the step that makes it work

Four k-means modes per side, not one Gaussian: sky is blue at the zenith,
white where the cloud is and pale at the horizon, and one blob over all
three rejects two of them.

Then modes holding under 3% of a side are discarded, and without that step
the whole thing fails on the case it was built for. A small flag deep in the
sky has both a high weight and a large distance from the boundary, so it
lands in the interior sample and teaches the model its own colour. It cannot
be excluded geometrically. It can be excluded by share.

Luminance is weighted at a quarter against chrominance for the same reason
the watershed's gradient is. Sky's variance is dominated by luminance, so at
equal weight the distribution is a long bright streak that a mid-grey flag
sits comfortably inside. A flag is separated by chrominance; a cloud is
separated by luminance alone. Not zero, or a dark bird against a bright sky
survives.

## Two tests, because either alone is wrong

Absolute — is this colour plausible under the category, as a chi-square on
the Mahalanobis distance. Comparative — is it likelier inside than outside.
A pixel must pass both.

The absolute test is what catches the flag, whose colour is far from *both*
sides and which the comparative test alone would leave at even odds. The
comparative test is what stops the absolute one needing a constant tuned per
category.

## What this cannot do, written down rather than left to be discovered

An intruder large enough to hold its own mode is kept. By share, a flag over
a fifth of the sky and a cloud bank over a fifth of the sky are the same
object, and colour does not separate them either — a white cloud is as far
from blue sky in chrominance as many intruders are.

So `min_cluster` is not a threshold with a correct value waiting to be
found; it is the trade-off itself, set where a photographic intruder falls.
Both ends are pinned by tests — `a_flag_in_the_sky_is_removed` and
`an_intruder_larger_than_min_cluster_survives` — so that moving the number
reads as moving the trade-off rather than as fixing a bug. The case left
open is a large unrecognised object in a clean category, which wants the
boundary snapped to watershed basins and is a different mechanism.

## Safe to apply without a control

It is subtractive: the output is the input times a factor in `0..=1`. The
worst failure available to it is losing part of a real sky, never gaining a
region, so a blue car below the horizon that was never in the mask cannot be
pulled into it. And a factor in `0..=1` cannot raise a sum, so `scene.rs`'s
partition still holds when every category is refined independently — the
weight taken off the flag lands in the unlisted remainder, which is where a
flag belongs, ADE20K having no class for one.

Every path without the evidence to judge returns the weights untouched and
says which path it took. A refinement that silently did nothing is
indistinguishable from the feature being off, and an empty seed set fitted
to a distribution would reject every pixel.

The signature is deliberately unchanged: categories are addressed by name,
not by index, so a sharper mask cannot create the stale-index hazard the
signature exists to guard against.

The example writes `<prefix>-<category>-refined.ppm` beside the coarse one,
never instead of it — whether this is an improvement is a comparative
judgement and one image cannot answer it.

Verified: fmt clean, clippy -D warnings clean, 57 dr-segment tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:16 +02:00
dtourolleandClaude Opus 5 8131706394 Regenerate the matrix over the merged rail work
The merge moved tagged lines in identity.rs, identity_ui.rs and the two
Slint files; the matrix tracks line numbers, so it goes stale on a move
alone. A merge commit does not run the pre-commit hook that would
normally have staged this.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:16 +02:00
dtourolleandClaude Opus 5 bfdb6d5e4b Cut the People rail's portraits off the blocking path
Opening Identity cut a portrait for every person before the screen was
allowed to appear. Measured against the reference library: 3.6 seconds,
of which 3.2 is 931 full 1024px proxy decodes — the fallback for faces
indexed before crops were stored beside them. The rest of the wait was
the rail building 14,268 rows, fixed in the commit before this.

`refresh` now draws with the portraits already cached and hands the rest
to `fill_covers`, which cuts them 8ms at a time — under half a frame —
behind a screen that is already up. Blocking work on that path goes from
4,965ms to 16ms on this library.

A timer rather than a thread. The work is a catalog query and a decode
against a `ThumbStore`, and both handles live on the UI thread; a worker
would need its own connection to the same file, which is what the sweep
and the regrouping pass do because they run for minutes and would
otherwise be unbounded. This is seconds of small, independent pieces, so
slicing answers the same question more cheaply.

`pending` is in rail order, and the rail is sorted by confirmed faces, so
the portraits the user is looking at are cut first. It patches single
rows rather than reloading: a reload would rebuild the model on every
tick, and a model replaced underneath the `ListView` is what the slicing
exists to avoid. Each patch checks the row still holds the person it was
started for — a stale index would draw a face beside somebody else's
name — and stops the fill when it does not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:16 +02:00
dtourolleandClaude Opus 5 d7b851f622 Build the People rail rows you can see, not the ones you cannot
`Flickable { VerticalLayout { for ... } }` instantiates every row. On the
reference library that is 14,268 row subtrees, which measures at 4.2
seconds before a pixel is drawn — and it was paid again on every action,
because every action reloads the model. Roughly half the wait between
clicking "Identity" and the screen appearing was this.

Slint's compiler has a virtualising path for a `for`; it is what makes
`std-widgets`' `ListView` cheap. It keys on the parent element's base
being *named* `ListView` and exposing the five lengths its layouting code
writes back, and a custom base is explicitly allowed. So `widgets.slint`
grows one, and `std-widgets` stays out of the file that establishes our
style. Measured on the real screen: 200,000 rail rows now render in 62ms.

Three things had to be true together, and each was silent on its own.

**The row height must be constant.** The rail hid a set-aside person with
`height: cond ? 52px : 0px`, a height that reads the model — so the
layout cannot place row N without building rows 0..N, and Slint builds
them all. The filtering moves to Rust, where the toggle was already
reloading anyway.

**The list must not be wrapped.** It carries its own stretch and preferred
size, having no natural height to offer; a Rectangle in between hands the
layout that Rectangle's constraints, which are taken from the list and
are therefore nothing.

**The panel must let it fill.** `Panel` lays its children out with
`alignment: start`, which gives each its preferred height — right for a
column of sliders, wrong for anything that scrolls. Hence `Panel.fill`.

Get any of them wrong and the rail renders empty, with no error and a
model full of people. All three were, in turn, before a headless render
of the real screen showed a blank rail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:16 +02:00
dtourolleandClaude Opus 5 f26f1ab694 Withhold the empty groups that fill the People rail
`load_people` returned every live person, and on the reference library
that was 14,268 rows of which 11,739 held no faces at all — 85 per cent
of the rail naming nobody and able to do nothing.

They are not a mystery. A regrouping pass creates a person per cluster;
the next pass moves those faces elsewhere and leaves the person it
emptied behind. `faces::prune_empty_unnamed` exists for exactly this and
runs only at the end of a pass, so nothing clears what accumulates
between them, and the Identity screen never prunes at all.

Each one cost a `for_person` query and a built row on every reload. This
withholds precisely the set the prune already treats as disposable —
empty, unnamed, not set aside — and no more.

Filtered rather than deleted: a screen is being drawn, not a catalog
repaired. Nothing is lost, a sync cannot resurrect what was never
removed, and the prune stays the one place that decides these can go.

An empty group with a *name* still shows. That one is not debris but the
symptom of a real failure — a named person whose faces were regrouped out
from under them — and hiding it would take away the only way to merge
them back.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:16 +02:00
dtourolleandClaude Opus 5 e8b072c815 Say the app is starting before there is anything to say it with
Nothing from our own code reached logcat on the tablet: 284 lines from
the app's PID during a launch, every one of them from Hwaps,
BlockMonitor, InputEventReceiver or nativeloader, and not even the
version banner android_main emits three statements in.

I could not find the fault in our wiring, and this commit does not claim
to fix it. What it does is make the next launch say which side is at
fault, and close a hole that is real whatever the answer turns out to be.

What reading rules out, so nobody repeats it: install does call
log::set_max_level. The Config carries an explicit tag and an explicit
max level, and its env_filter is None, so android_logger's enabled() and
filter_matches() both pass an Info record. Tee::enabled delegates to the
console and gates nothing else. set_boxed_logger cannot have failed —
its Err path drops the Tee, taking the LogFile with it, and the file on
the device has content. And Tee::log calls console.log() unconditionally
*before* the file write, which is itself gated on console.enabled(), so
every line that reached the file proves the AndroidLogger was handed the
same record. Nothing else in the graph installs a logger; android-activity
and Slint's backend do not. The diff that introduced this changed the
level, the tag and the Config not at all — init_once and set_boxed_logger
leave the same logger installed at the same level.

That leaves below __android_log_write, which no amount of reading this
file can reach.

So: the console logger is now built first, and one line goes through it
directly, before the state directory and before the log file. Two things
follow. Everything between android_main's first statement and install
returning — external_data_path, create_dir_all and an open on a FUSE
volume the system may still be mounting — currently has no surface at
all to fail on; that window is what logcat is for, and it now has a line
in it. And when the log is silent, that line separates the two cases:
present with the log::info! lines below it missing is the facade, absent
along with them is liblog not delivering this process's records.

It goes through Log::log rather than log::info!, which is not a style
choice: the facade's maximum level is Off until install sets it, so a
log::info! there compiles and emits nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 17:31:41 +02:00
dtourolleandClaude Opus 5 c62edd3317 Give the log a mode that the way off the device can open
The log file has been on external storage since it existed, and the
reason is stated at length in two places: /data/data/<pkg>/files needs
run-as against a debuggable build, /sdcard/Android/data/<pkg>/files is a
plain adb pull from any build, and a log nobody can retrieve is not a
diagnostic. The file was then opened 0600, which cancels that decision
out. On the tablet:

    adb pull      -> remote open failed: Permission denied
    adb shell cat -> Permission denied
    run-as        -> package not debuggable

Every route off the device closed at once, on a file whose whole purpose
is to leave the device.

The mode is now per platform, because "who may read this" has two
different answers and the directory above the file is what makes them
differ. On the desktop, 0600 as before: $XDG_STATE_HOME/darkroom is in a
home directory on a machine that may have other accounts, and nothing
about that directory stops another local user reading a world-readable
file. On Android, 0644: /sdcard/Android/data is drwxrws--x
media_rw:ext_data_rw, so no other app can enter this app's subdirectory
and anyone who can traverse it is holding the unlocked tablet, which
already gets them the photographs the log merely names. What the read
bits buy is adb pull, which runs as shell — able to traverse a --x
directory, but then obliged to open the file as other.

The mode is also applied twice, and the second one is the fix rather
than belt and braces. OpenOptions::mode is a request: the kernel ANDs it
with the process umask, and an Android application process inherits
0o077 from the zygote, so asking for 0644 there creates 0600 and reports
nothing. It is ignored outright on a file that already exists, which
every launch after the first has. fchmod is subject to neither, and is
what the second call makes.

The comment claiming the mode was "ignored by the FAT-derived filesystem
Android presents as external storage" is gone with it. The device says
otherwise: the file it produced was 0600 exactly.

Two tests. One pins the literal mode per platform — only the desktop arm
can run under cargo test, and the comment says so rather than implying
the Android number is covered. The other reopens a log left behind with
the wrong mode, which is the one assertion on the host that fails if the
fchmod is deleted, since OpenOptions::mode cannot touch a file that is
already there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 17:31:10 +02:00
264 changed files with 58009 additions and 3922 deletions
+10
View File
@@ -10,3 +10,13 @@
# detects exactly that and fails with an instruction rather than embedding the
# pointer and failing at inference time.
*.onnx filter=lfs diff=lfs merge=lfs -text
# Test photographs live in LFS too, and are fetched only by the tests that
# need them.
#
# `fixtures/**` holds real camera files — a twelve-frame panorama set is
# 325 MB — and CI's `git lfs pull` excludes the directory, so a checkout
# carries pointers there until a merge test asks for the frames. Same
# reasoning as the models, with the opposite default: the model is not
# optional and the fixtures are.
fixtures/** filter=lfs diff=lfs merge=lfs -text
+1 -1
View File
@@ -148,7 +148,7 @@ jobs:
| while read -r key; do git config --local --unset-all "$key"; done || true
git config --local lfs.url \
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
git lfs pull
git lfs pull --exclude="fixtures/**"
- name: Cache cargo
uses: actions/cache@v4
+108 -2
View File
@@ -96,7 +96,7 @@ jobs:
| while read -r key; do git config --local --unset-all "$key"; done || true
git config --local lfs.url \
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
git lfs pull
git lfs pull --exclude="fixtures/**"
ls -lR models/
- name: Cache cargo
@@ -213,7 +213,7 @@ jobs:
| while read -r key; do git config --local --unset-all "$key"; done || true
git config --local lfs.url \
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
git lfs pull
git lfs pull --exclude="fixtures/**"
ls -lR models/
- name: Cache cargo
@@ -365,6 +365,112 @@ jobs:
path: target-android/apk/darkroom.apk
if-no-files-found: error
windows-image:
uses: ./.gitea/workflows/windows-image.yml
# TRACES: FR-PLAT-WIN-3
# The Windows executable and its installer, cross-built from Linux
# (docs/windows.md §7). No Windows machine anywhere in this job: what it
# can prove is that the binary links, is a Windows executable with no
# MinGW runtime imports, starts under Wine, and that the installer installs
# and uninstalls under Wine. What it cannot prove — a Vulkan device, a
# render, the secret store — is a release step on a real machine (§6).
windows:
runs-on: linux/amd64
name: Windows (x86_64, cross)
needs: windows-image
container:
image: gitea.tourolle.paris/dtourolle/darkroom-windows:latest
env:
CARGO_INCREMENTAL: 0
CARGO_PROFILE_DEV_DEBUG: 0
CARGO_TARGET_DIR: target-windows
# Wine keeps its prefix under $HOME, which the image points at a
# directory that does not exist in a fresh container.
HOME: /tmp/home
steps:
- name: Checkout
uses: actions/checkout@v4
# Same step as the desktop leg: the models are LFS objects and the
# packager refuses pointers.
- name: Fetch the models
env:
LFS_TOKEN: ${{ secrets.GITEA_TOKEN || github.token }}
run: |
set -e
git lfs install --local
git config --local --get-regexp '^http\..*extraheader$' \
| cut -d' ' -f1 | sort -u \
| while read -r key; do git config --local --unset-all "$key"; done || true
git config --local lfs.url \
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
git lfs pull --exclude="fixtures/**"
ls -l models/face models/scene
- name: Cache cargo
uses: actions/cache@v4
with:
path: |
/opt/cargo/registry
target-windows
key: windows-${{ hashFiles('**/Cargo.lock') }}
# The cfg(windows) branches are linted here and nowhere else: the
# desktop leg's clippy never compiles them.
- name: Clippy for the target
run: cargo clippy --release --target x86_64-pc-windows-gnu -p darkroom-desktop -- -D warnings
- name: Build
run: cargo build --release --target x86_64-pc-windows-gnu -p darkroom-desktop
- name: Smoke-test the executable
run: |
set -e
mkdir -p "$HOME"
EXE=target-windows/x86_64-pc-windows-gnu/release/darkroom-desktop.exe
file "$EXE"
file "$EXE" | grep -q 'PE32+' || { echo "FAIL: not a PE32+ executable"; exit 1; }
file "$EXE" | grep -q '(GUI)' || { echo "FAIL: not a GUI-subsystem executable"; exit 1; }
if x86_64-w64-mingw32-objdump -p "$EXE" | grep -iE 'libwinpthread|libgcc|libstdc'; then
echo "FAIL: the executable imports a MinGW runtime DLL"
exit 1
fi
x86_64-w64-mingw32-objdump -p "$EXE" | grep 'DLL Name' | sort -u
wineboot --init >/dev/null 2>&1 || true
OUT=$(wine "$EXE" --version 2>/dev/null)
echo "wine: $OUT"
echo "$OUT" | grep -q '^darkroom-desktop ' || { echo "FAIL: --version did not answer under Wine"; exit 1; }
- name: Package the installer
run: bash docker/windows/package.sh
- name: Smoke-test the installer
run: |
set -e
SETUP=$(ls target-windows/installer/DarkRoom-*-x86_64-setup.exe)
file "$SETUP" | grep -q 'PE32+' || { echo "FAIL: the installer is not 64-bit"; exit 1; }
wine "$SETUP" /S 2>/dev/null
INST=$(echo "$HOME"/.wine/drive_c/users/*/AppData/Local/Programs/DarkRoom)
ls "$INST"
[ "$(ls "$INST/models" | wc -l)" = 7 ] || { echo "FAIL: expected 7 model files"; exit 1; }
wine reg query 'HKCU\Software\Microsoft\Windows\CurrentVersion\Uninstall\DarkRoom' 2>/dev/null \
| grep -q DisplayVersion || { echo "FAIL: no uninstall registry key"; exit 1; }
wine "$INST/darkroom.exe" --version 2>/dev/null | grep -q '^darkroom-desktop ' \
|| { echo "FAIL: the installed executable does not run"; exit 1; }
wine "$INST/uninstall.exe" /S 2>/dev/null
sleep 3
[ ! -e "$INST" ] || { echo "FAIL: uninstall left $INST behind"; ls -R "$INST"; exit 1; }
echo "OK: installed and uninstalled under Wine"
- name: Upload the installer
uses: actions/upload-artifact@v3
with:
name: darkroom-windows-x86_64-setup
path: target-windows/installer/DarkRoom-*-x86_64-setup.exe
if-no-files-found: error
layering:
runs-on: linux/amd64
name: Layer separation
+170
View File
@@ -0,0 +1,170 @@
name: '🐳 Windows image'
# Builds and pushes gitea.tourolle.paris/dtourolle/darkroom-windows, the job
# container for the Windows leg of build-and-test.yml.
#
# The same shape as android-image.yml, for the same reason that one exists:
# an image that lives only on a developer's laptop is a job that dies at
# `docker pull`. Built from docker/windows, tagged by that directory's tree
# id, skipped when the registry already has it.
#
# Called by build-and-test.yml on every push, and runnable by hand via
# workflow_dispatch. It is cheap when nothing changed — see the guard below.
on:
workflow_call:
inputs:
force:
description: 'Rebuild even if the registry already has this image ("true"/"false")'
type: string
default: 'false'
workflow_dispatch:
inputs:
force:
description: 'Rebuild even if the registry already has this image ("true"/"false")'
type: string
default: 'false'
# Gitea's act_runner mangles boolean workflow inputs passed through an
# expression — they arrive as false regardless of what was sent. Every input
# here is a string compared with == 'true', as in KPN's docker.yaml.
env:
IMAGE: gitea.tourolle.paris/dtourolle/darkroom-windows
jobs:
build:
runs-on: linux/amd64
name: Build and push
# Deliberately NOT in a container: this job needs the host Docker daemon to
# build an image, and the host's cached ~/.docker/config.json to push it.
# That is also why there is no `docker login` step — the runner host was
# authenticated to the registry during setup.
steps:
# The host has no Node, so the JS-based actions/checkout cannot run here.
# A minimal shallow fetch with plain git gets the same tree.
- name: Checkout
run: |
set -e
git init -q .
git remote add origin "${{ github.server_url }}/${{ github.repository }}.git"
git -c http.extraheader="AUTHORIZATION: basic $(printf '%s' '${{ github.actor }}:${{ github.token }}' | base64 -w0)" \
fetch --depth 1 origin "${{ github.sha }}"
git checkout -q FETCH_HEAD
# The image is tagged by the content of docker/windows, not by the commit
# that happened to touch it. `git rev-parse HEAD:<dir>` is the tree object
# id — it changes when and only when a file in that directory changes, so
# an unrelated push reuses the existing image and a Dockerfile edit can
# never silently keep serving a stale `latest`.
#
# Using the commit sha instead would rebuild 2.5 GB on every push; using a
# paths-filter action would need a container that has Node, and the only
# one this repo would reach for is the very image being built.
- name: Resolve image tag
id: tag
run: |
set -e
TREE=$(git rev-parse HEAD:docker/windows)
echo "tree=$TREE" >> "$GITHUB_OUTPUT"
echo "docker/windows tree: $TREE"
# Skip the build when the registry already holds this exact content. This
# is what keeps the job a few seconds long on a normal push, and what
# makes it self-healing: if the tag is missing for any reason, including
# the image having never been pushed at all, it gets built here.
#
# The probe is curl against the registry API, NOT `docker manifest
# inspect`. The latter exits 1 on this registry even for tags that are
# demonstrably present — jellytau-builder:latest answers HTTP 200 to the
# API while `docker manifest inspect` reports "manifest unknown" for it.
# Trusting that would have rebuilt 7 GB on every single push.
#
# A HEAD request also gives the digest for free, which is how the repoint
# decision below is made without pulling any layers.
- name: Query registry
id: check
env:
# The runner's own credentials, so this does not depend on how the
# host's ~/.docker/config.json happens to be set up.
REG_USER: ${{ github.actor }}
REG_PASS: ${{ github.token }}
TREE: ${{ steps.tag.outputs.tree }}
run: |
set -eu
ACCEPT='application/vnd.oci.image.index.v1+json,application/vnd.docker.distribution.manifest.v2+json,application/vnd.oci.image.manifest.v1+json,application/vnd.docker.distribution.manifest.list.v2+json'
API="https://gitea.tourolle.paris/v2/dtourolle/darkroom-windows/manifests"
# Prints "<http-status> <digest-or-empty>" for a tag.
probe() {
curl -sI -u "$REG_USER:$REG_PASS" -H "Accept: $ACCEPT" "$API/$1" \
| tr -d '\r' \
| awk 'BEGIN{s="000";d=""} /^HTTP/{s=$2} tolower($1)=="docker-content-digest:"{d=$2} END{print s, d}'
}
read -r TREE_STATUS TREE_DIGEST <<EOF
$(probe "$TREE")
EOF
read -r LATEST_STATUS LATEST_DIGEST <<EOF
$(probe latest)
EOF
echo "tag $TREE -> HTTP $TREE_STATUS ${TREE_DIGEST:-(no digest)}"
echo "tag latest -> HTTP $LATEST_STATUS ${LATEST_DIGEST:-(no digest)}"
# Build unless the registry definitively confirms this content is
# already there. An auth failure or an unreachable registry lands
# here too, and rebuilding needlessly is the safe direction to fail —
# skipping a build that was needed is what breaks the Windows job.
if [ "${{ inputs.force }}" = "true" ]; then
echo "forced rebuild requested"
echo "build=true" >> "$GITHUB_OUTPUT"
echo "repoint=false" >> "$GITHUB_OUTPUT"
elif [ "$TREE_STATUS" != "200" ]; then
echo "registry does not have this content — building"
echo "build=true" >> "$GITHUB_OUTPUT"
echo "repoint=false" >> "$GITHUB_OUTPUT"
elif [ -n "$TREE_DIGEST" ] && [ "$TREE_DIGEST" = "$LATEST_DIGEST" ]; then
echo "registry is already correct — nothing to do"
echo "build=false" >> "$GITHUB_OUTPUT"
echo "repoint=false" >> "$GITHUB_OUTPUT"
else
echo "content is present but latest points elsewhere — repointing"
echo "build=false" >> "$GITHUB_OUTPUT"
echo "repoint=true" >> "$GITHUB_OUTPUT"
fi
# Context is docker/windows, matching the README's build command. The
# Dockerfile COPYs nothing from the repo, so it needs no wider context —
# and a narrow context keeps the daemon from tarring up the whole tree,
# target/ included.
- name: Build
if: ${{ steps.check.outputs.build == 'true' }}
run: |
set -e
docker build \
-t "$IMAGE:${{ steps.tag.outputs.tree }}" \
-t "$IMAGE:latest" \
docker/windows
# Both tags are pushed: the tree tag is what the guard above looks for on
# the next run, and `latest` is what build-and-test.yml pulls.
- name: Push
if: ${{ steps.check.outputs.build == 'true' }}
run: |
set -e
docker push "$IMAGE:${{ steps.tag.outputs.tree }}"
docker push "$IMAGE:latest"
# A cache hit on the tree tag says nothing about where `latest` points — a
# reverted Dockerfile or a build from another branch can leave it on
# different content. This runs only when the digests above actually
# disagree, so the common case costs nothing; the layers are already in
# the registry, so the push that follows uploads a manifest, not 2.5 GB.
- name: Repoint latest
if: ${{ steps.check.outputs.repoint == 'true' }}
run: |
set -e
docker pull "$IMAGE:${{ steps.tag.outputs.tree }}"
docker tag "$IMAGE:${{ steps.tag.outputs.tree }}" "$IMAGE:latest"
docker push "$IMAGE:latest"
+1
View File
@@ -23,3 +23,4 @@ tools/film-profiles/upstream/
# checkout, so it is larger than the repository it sits in.
/.flatpak-builder/
/build/
__pycache__/
Generated
+79 -24
View File
@@ -1221,7 +1221,7 @@ checksum = "f27ae1dd37df86211c42e150270f82743308803d90a6f6e6651cd730d5e1732f"
[[package]]
name = "darkroom-android"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"android_logger",
"dr-plat",
@@ -1234,13 +1234,14 @@ dependencies = [
[[package]]
name = "darkroom-desktop"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"anyhow",
"dr-plat",
"dr-ui",
"env_logger",
"log",
"winresource",
]
[[package]]
@@ -1407,7 +1408,7 @@ checksum = "d8b14ccef22fc6f5a8f4d7d768562a182c04ce9a3b3157b91390b52ddfdf1a76"
[[package]]
name = "dr-bench"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"anyhow",
"dr-catalog",
@@ -1424,7 +1425,7 @@ dependencies = [
[[package]]
name = "dr-catalog"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"dr-face",
"dr-plat",
@@ -1439,7 +1440,7 @@ dependencies = [
[[package]]
name = "dr-decode"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"dr-types",
"env_logger",
@@ -1453,7 +1454,7 @@ dependencies = [
[[package]]
name = "dr-export"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"dr-decode",
"dr-gpu",
@@ -1464,6 +1465,7 @@ dependencies = [
"log",
"png",
"pollster",
"rawler",
"thiserror 2.0.20",
"tiff",
"zune-jpeg 0.4.21",
@@ -1471,20 +1473,20 @@ dependencies = [
[[package]]
name = "dr-face"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"dr-inference-engine",
"env_logger",
"log",
"ndarray",
"ort",
"ort-tract",
"thiserror 2.0.20",
"zune-jpeg 0.4.21",
]
[[package]]
name = "dr-film"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"log",
"serde",
@@ -1493,11 +1495,12 @@ dependencies = [
[[package]]
name = "dr-gpu"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"bytemuck",
"dr-decode",
"dr-film",
"dr-pano",
"dr-pipeline",
"dr-segment",
"dr-types",
@@ -1508,9 +1511,23 @@ dependencies = [
"wgpu",
]
[[package]]
name = "dr-inference-engine"
version = "0.13.1"
dependencies = [
"libloading",
"log",
"ort",
"ort-sys",
"ort-tract",
"serde",
"serde_json",
"thiserror 2.0.20",
]
[[package]]
name = "dr-ingest"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"dr-plat",
"dr-types",
@@ -1522,15 +1539,29 @@ dependencies = [
[[package]]
name = "dr-lens"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"lensfun",
"log",
]
[[package]]
name = "dr-pano"
version = "0.13.1"
dependencies = [
"dr-decode",
"dr-inference-engine",
"dr-types",
"env_logger",
"log",
"ndarray",
"ort",
"thiserror 2.0.20",
]
[[package]]
name = "dr-pipeline"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"dr-types",
"log",
@@ -1539,7 +1570,7 @@ dependencies = [
[[package]]
name = "dr-plat"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"android-native-keyring-store",
"dr-types",
@@ -1555,7 +1586,7 @@ dependencies = [
[[package]]
name = "dr-preset-xmp"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"dr-pipeline",
"log",
@@ -1565,20 +1596,20 @@ dependencies = [
[[package]]
name = "dr-segment"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"dr-inference-engine",
"env_logger",
"log",
"ndarray",
"ort",
"ort-tract",
"thiserror 2.0.20",
"zune-jpeg 0.4.21",
]
[[package]]
name = "dr-sync"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"async-trait",
"dr-plat",
@@ -1592,7 +1623,7 @@ dependencies = [
[[package]]
name = "dr-sync-folder"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"async-trait",
"dr-sync",
@@ -1604,7 +1635,7 @@ dependencies = [
[[package]]
name = "dr-sync-nextcloud"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"async-trait",
"dr-decode",
@@ -1626,7 +1657,7 @@ dependencies = [
[[package]]
name = "dr-thumbs"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"dr-types",
"jpeg-encoder",
@@ -1638,7 +1669,7 @@ dependencies = [
[[package]]
name = "dr-types"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"serde",
"serde_json",
@@ -1647,7 +1678,7 @@ dependencies = [
[[package]]
name = "dr-ui"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"anyhow",
"async-trait",
@@ -1657,7 +1688,10 @@ dependencies = [
"dr-face",
"dr-film",
"dr-gpu",
"dr-inference-engine",
"dr-ingest",
"dr-lens",
"dr-pano",
"dr-pipeline",
"dr-plat",
"dr-preset-xmp",
@@ -1667,6 +1701,7 @@ dependencies = [
"dr-sync-nextcloud",
"dr-thumbs",
"dr-types",
"dr-xmp",
"env_logger",
"jni 0.22.4",
"log",
@@ -1683,6 +1718,16 @@ dependencies = [
"wgpu",
]
[[package]]
name = "dr-xmp"
version = "0.13.1"
dependencies = [
"dr-types",
"log",
"quick-xml",
"thiserror 2.0.20",
]
[[package]]
name = "drm"
version = "0.14.1"
@@ -6976,7 +7021,7 @@ checksum = "8df9b6e13f2d32c91b9bd719c00d1958837bc7dec474d94952798cc8e69eeec3"
[[package]]
name = "traceability"
version = "0.10.0"
version = "0.13.1"
dependencies = [
"anyhow",
"serde",
@@ -8376,6 +8421,16 @@ dependencies = [
"memchr",
]
[[package]]
name = "winresource"
version = "0.1.31"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0986a8b1d586b7d3e4fe3d9ea39fb451ae22869dcea4aa109d287a374d866087"
dependencies = [
"toml 1.1.4+spec-1.1.0",
"version_check",
]
[[package]]
name = "wit-bindgen"
version = "0.57.1"
+27 -6
View File
@@ -8,15 +8,18 @@ members = [
"core/dr-export",
"core/dr-face",
"core/dr-film",
"core/dr-inference-engine",
"core/dr-ingest",
"core/dr-gpu",
"core/dr-lens",
"core/dr-pano",
"core/dr-pipeline",
"core/dr-preset-xmp",
"core/dr-segment",
"core/dr-sync",
"core/dr-sync-folder",
"core/dr-sync-nextcloud",
"core/dr-xmp",
"platform/dr-plat",
"ui/dr-ui",
"apps/darkroom-desktop",
@@ -26,7 +29,7 @@ members = [
]
[workspace.package]
version = "0.10.0"
version = "0.13.1"
edition = "2021"
rust-version = "1.92"
license = "GPL-3.0-or-later"
@@ -44,9 +47,14 @@ dr-export = { path = "core/dr-export" }
# `features = ["inference"]`.
dr-face = { path = "core/dr-face", default-features = false }
dr-film = { path = "core/dr-film" }
# `tract` on by default so a test binary can open a session with nothing
# installed; the apps add `native` to look for a runtime file (docs/inference.md §3).
dr-inference-engine = { path = "core/dr-inference-engine" }
dr-ingest = { path = "core/dr-ingest" }
dr-gpu = { path = "core/dr-gpu" }
dr-lens = { path = "core/dr-lens" }
# Optional runtime, like `dr-segment`: the geometry never needs a model.
dr-pano = { path = "core/dr-pano", default-features = false }
dr-pipeline = { path = "core/dr-pipeline" }
dr-preset-xmp = { path = "core/dr-preset-xmp" }
# `default-features = false` belongs *here*, not on each dependant: a member
@@ -59,6 +67,7 @@ dr-plat = { path = "platform/dr-plat" }
dr-sync = { path = "core/dr-sync" }
dr-sync-folder = { path = "core/dr-sync-folder" }
dr-sync-nextcloud = { path = "core/dr-sync-nextcloud" }
dr-xmp = { path = "core/dr-xmp" }
dr-ui = { path = "ui/dr-ui" }
# GPU + UI
@@ -231,13 +240,25 @@ ort-tract = "0.4"
ndarray = "0.17"
[profile.dev]
# Dependencies optimised even in dev builds — wgpu and image decoding are
# unusably slow otherwise, and they rarely need debugging.
# Dev builds are tuned for how fast they *compile*, not for how fast they run.
# Optimisation is a release concern; `[profile.release]` below is where it
# belongs.
#
# This deliberately reverses an earlier choice. Dependencies used to be built
# at `opt-level = 2` here, because wgpu and image decoding are slow without it.
# That is still true, and it is the price: a debug run of the app, and the
# decode- and GPU-heavy tests, are slower than they were. What it buys is that
# nothing has to be optimised before it can be compiled — which is the cost
# paid on every edit, by every worktree, rather than only when something is
# actually run.
#
# If a particular crate turns out to be the one that makes a test unbearable,
# raise it alone rather than restoring the blanket rule:
#
# [profile.dev.package.zune-jpeg]
# opt-level = 2
opt-level = 0
[profile.dev.package."*"]
opt-level = 2
[profile.release]
lto = "thin"
codegen-units = 1
+232
View File
@@ -0,0 +1,232 @@
GNU GENERAL PUBLIC LICENSE
Version 3, 29 June 2007
Copyright © 2007 Free Software Foundation, Inc. <https://fsf.org/>
Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed.
Preamble
The GNU General Public License is a free, copyleft license for software and other kinds of works.
The licenses for most software and other practical works are designed to take away your freedom to share and change the works. By contrast, the GNU General Public License is intended to guarantee your freedom to share and change all versions of a program--to make sure it remains free software for all its users. We, the Free Software Foundation, use the GNU General Public License for most of our software; it applies also to any other work released this way by its authors. You can apply it to your programs, too.
When we speak of free software, we are referring to freedom, not price. Our General Public Licenses are designed to make sure that you have the freedom to distribute copies of free software (and charge for them if you wish), that you receive source code or can get it if you want it, that you can change the software or use pieces of it in new free programs, and that you know you can do these things.
To protect your rights, we need to prevent others from denying you these rights or asking you to surrender the rights. Therefore, you have certain responsibilities if you distribute copies of the software, or if you modify it: responsibilities to respect the freedom of others.
For example, if you distribute copies of such a program, whether gratis or for a fee, you must pass on to the recipients the same freedoms that you received. You must make sure that they, too, receive or can get the source code. And you must show them these terms so they know their rights.
Developers that use the GNU GPL protect your rights with two steps: (1) assert copyright on the software, and (2) offer you this License giving you legal permission to copy, distribute and/or modify it.
For the developers' and authors' protection, the GPL clearly explains that there is no warranty for this free software. For both users' and authors' sake, the GPL requires that modified versions be marked as changed, so that their problems will not be attributed erroneously to authors of previous versions.
Some devices are designed to deny users access to install or run modified versions of the software inside them, although the manufacturer can do so. This is fundamentally incompatible with the aim of protecting users' freedom to change the software. The systematic pattern of such abuse occurs in the area of products for individuals to use, which is precisely where it is most unacceptable. Therefore, we have designed this version of the GPL to prohibit the practice for those products. If such problems arise substantially in other domains, we stand ready to extend this provision to those domains in future versions of the GPL, as needed to protect the freedom of users.
Finally, every program is threatened constantly by software patents. States should not allow patents to restrict development and use of software on general-purpose computers, but in those that do, we wish to avoid the special danger that patents applied to a free program could make it effectively proprietary. To prevent this, the GPL assures that patents cannot be used to render the program non-free.
The precise terms and conditions for copying, distribution and modification follow.
TERMS AND CONDITIONS
0. Definitions.
“This License” refers to version 3 of the GNU General Public License.
“Copyright” also means copyright-like laws that apply to other kinds of works, such as semiconductor masks.
“The Program” refers to any copyrightable work licensed under this License. Each licensee is addressed as “you”. “Licensees” and “recipients” may be individuals or organizations.
To “modify” a work means to copy from or adapt all or part of the work in a fashion requiring copyright permission, other than the making of an exact copy. The resulting work is called a “modified version” of the earlier work or a work “based on” the earlier work.
A “covered work” means either the unmodified Program or a work based on the Program.
To “propagate” a work means to do anything with it that, without permission, would make you directly or secondarily liable for infringement under applicable copyright law, except executing it on a computer or modifying a private copy. Propagation includes copying, distribution (with or without modification), making available to the public, and in some countries other activities as well.
To “convey” a work means any kind of propagation that enables other parties to make or receive copies. Mere interaction with a user through a computer network, with no transfer of a copy, is not conveying.
An interactive user interface displays “Appropriate Legal Notices” to the extent that it includes a convenient and prominently visible feature that (1) displays an appropriate copyright notice, and (2) tells the user that there is no warranty for the work (except to the extent that warranties are provided), that licensees may convey the work under this License, and how to view a copy of this License. If the interface presents a list of user commands or options, such as a menu, a prominent item in the list meets this criterion.
1. Source Code.
The “source code” for a work means the preferred form of the work for making modifications to it. “Object code” means any non-source form of a work.
A “Standard Interface” means an interface that either is an official standard defined by a recognized standards body, or, in the case of interfaces specified for a particular programming language, one that is widely used among developers working in that language.
The “System Libraries” of an executable work include anything, other than the work as a whole, that (a) is included in the normal form of packaging a Major Component, but which is not part of that Major Component, and (b) serves only to enable use of the work with that Major Component, or to implement a Standard Interface for which an implementation is available to the public in source code form. A “Major Component”, in this context, means a major essential component (kernel, window system, and so on) of the specific operating system (if any) on which the executable work runs, or a compiler used to produce the work, or an object code interpreter used to run it.
The “Corresponding Source” for a work in object code form means all the source code needed to generate, install, and (for an executable work) run the object code and to modify the work, including scripts to control those activities. However, it does not include the work's System Libraries, or general-purpose tools or generally available free programs which are used unmodified in performing those activities but which are not part of the work. For example, Corresponding Source includes interface definition files associated with source files for the work, and the source code for shared libraries and dynamically linked subprograms that the work is specifically designed to require, such as by intimate data communication or control flow between those subprograms and other parts of the work.
The Corresponding Source need not include anything that users can regenerate automatically from other parts of the Corresponding Source.
The Corresponding Source for a work in source code form is that same work.
2. Basic Permissions.
All rights granted under this License are granted for the term of copyright on the Program, and are irrevocable provided the stated conditions are met. This License explicitly affirms your unlimited permission to run the unmodified Program. The output from running a covered work is covered by this License only if the output, given its content, constitutes a covered work. This License acknowledges your rights of fair use or other equivalent, as provided by copyright law.
You may make, run and propagate covered works that you do not convey, without conditions so long as your license otherwise remains in force. You may convey covered works to others for the sole purpose of having them make modifications exclusively for you, or provide you with facilities for running those works, provided that you comply with the terms of this License in conveying all material for which you do not control copyright. Those thus making or running the covered works for you must do so exclusively on your behalf, under your direction and control, on terms that prohibit them from making any copies of your copyrighted material outside their relationship with you.
Conveying under any other circumstances is permitted solely under the conditions stated below. Sublicensing is not allowed; section 10 makes it unnecessary.
3. Protecting Users' Legal Rights From Anti-Circumvention Law.
No covered work shall be deemed part of an effective technological measure under any applicable law fulfilling obligations under article 11 of the WIPO copyright treaty adopted on 20 December 1996, or similar laws prohibiting or restricting circumvention of such measures.
When you convey a covered work, you waive any legal power to forbid circumvention of technological measures to the extent such circumvention is effected by exercising rights under this License with respect to the covered work, and you disclaim any intention to limit operation or modification of the work as a means of enforcing, against the work's users, your or third parties' legal rights to forbid circumvention of technological measures.
4. Conveying Verbatim Copies.
You may convey verbatim copies of the Program's source code as you receive it, in any medium, provided that you conspicuously and appropriately publish on each copy an appropriate copyright notice; keep intact all notices stating that this License and any non-permissive terms added in accord with section 7 apply to the code; keep intact all notices of the absence of any warranty; and give all recipients a copy of this License along with the Program.
You may charge any price or no price for each copy that you convey, and you may offer support or warranty protection for a fee.
5. Conveying Modified Source Versions.
You may convey a work based on the Program, or the modifications to produce it from the Program, in the form of source code under the terms of section 4, provided that you also meet all of these conditions:
a) The work must carry prominent notices stating that you modified it, and giving a relevant date.
b) The work must carry prominent notices stating that it is released under this License and any conditions added under section 7. This requirement modifies the requirement in section 4 to “keep intact all notices”.
c) You must license the entire work, as a whole, under this License to anyone who comes into possession of a copy. This License will therefore apply, along with any applicable section 7 additional terms, to the whole of the work, and all its parts, regardless of how they are packaged. This License gives no permission to license the work in any other way, but it does not invalidate such permission if you have separately received it.
d) If the work has interactive user interfaces, each must display Appropriate Legal Notices; however, if the Program has interactive interfaces that do not display Appropriate Legal Notices, your work need not make them do so.
A compilation of a covered work with other separate and independent works, which are not by their nature extensions of the covered work, and which are not combined with it such as to form a larger program, in or on a volume of a storage or distribution medium, is called an “aggregate” if the compilation and its resulting copyright are not used to limit the access or legal rights of the compilation's users beyond what the individual works permit. Inclusion of a covered work in an aggregate does not cause this License to apply to the other parts of the aggregate.
6. Conveying Non-Source Forms.
You may convey a covered work in object code form under the terms of sections 4 and 5, provided that you also convey the machine-readable Corresponding Source under the terms of this License, in one of these ways:
a) Convey the object code in, or embodied in, a physical product (including a physical distribution medium), accompanied by the Corresponding Source fixed on a durable physical medium customarily used for software interchange.
b) Convey the object code in, or embodied in, a physical product (including a physical distribution medium), accompanied by a written offer, valid for at least three years and valid for as long as you offer spare parts or customer support for that product model, to give anyone who possesses the object code either (1) a copy of the Corresponding Source for all the software in the product that is covered by this License, on a durable physical medium customarily used for software interchange, for a price no more than your reasonable cost of physically performing this conveying of source, or (2) access to copy the Corresponding Source from a network server at no charge.
c) Convey individual copies of the object code with a copy of the written offer to provide the Corresponding Source. This alternative is allowed only occasionally and noncommercially, and only if you received the object code with such an offer, in accord with subsection 6b.
d) Convey the object code by offering access from a designated place (gratis or for a charge), and offer equivalent access to the Corresponding Source in the same way through the same place at no further charge. You need not require recipients to copy the Corresponding Source along with the object code. If the place to copy the object code is a network server, the Corresponding Source may be on a different server (operated by you or a third party) that supports equivalent copying facilities, provided you maintain clear directions next to the object code saying where to find the Corresponding Source. Regardless of what server hosts the Corresponding Source, you remain obligated to ensure that it is available for as long as needed to satisfy these requirements.
e) Convey the object code using peer-to-peer transmission, provided you inform other peers where the object code and Corresponding Source of the work are being offered to the general public at no charge under subsection 6d.
A separable portion of the object code, whose source code is excluded from the Corresponding Source as a System Library, need not be included in conveying the object code work.
A “User Product” is either (1) a “consumer product”, which means any tangible personal property which is normally used for personal, family, or household purposes, or (2) anything designed or sold for incorporation into a dwelling. In determining whether a product is a consumer product, doubtful cases shall be resolved in favor of coverage. For a particular product received by a particular user, “normally used” refers to a typical or common use of that class of product, regardless of the status of the particular user or of the way in which the particular user actually uses, or expects or is expected to use, the product. A product is a consumer product regardless of whether the product has substantial commercial, industrial or non-consumer uses, unless such uses represent the only significant mode of use of the product.
“Installation Information” for a User Product means any methods, procedures, authorization keys, or other information required to install and execute modified versions of a covered work in that User Product from a modified version of its Corresponding Source. The information must suffice to ensure that the continued functioning of the modified object code is in no case prevented or interfered with solely because modification has been made.
If you convey an object code work under this section in, or with, or specifically for use in, a User Product, and the conveying occurs as part of a transaction in which the right of possession and use of the User Product is transferred to the recipient in perpetuity or for a fixed term (regardless of how the transaction is characterized), the Corresponding Source conveyed under this section must be accompanied by the Installation Information. But this requirement does not apply if neither you nor any third party retains the ability to install modified object code on the User Product (for example, the work has been installed in ROM).
The requirement to provide Installation Information does not include a requirement to continue to provide support service, warranty, or updates for a work that has been modified or installed by the recipient, or for the User Product in which it has been modified or installed. Access to a network may be denied when the modification itself materially and adversely affects the operation of the network or violates the rules and protocols for communication across the network.
Corresponding Source conveyed, and Installation Information provided, in accord with this section must be in a format that is publicly documented (and with an implementation available to the public in source code form), and must require no special password or key for unpacking, reading or copying.
7. Additional Terms.
“Additional permissions” are terms that supplement the terms of this License by making exceptions from one or more of its conditions. Additional permissions that are applicable to the entire Program shall be treated as though they were included in this License, to the extent that they are valid under applicable law. If additional permissions apply only to part of the Program, that part may be used separately under those permissions, but the entire Program remains governed by this License without regard to the additional permissions.
When you convey a copy of a covered work, you may at your option remove any additional permissions from that copy, or from any part of it. (Additional permissions may be written to require their own removal in certain cases when you modify the work.) You may place additional permissions on material, added by you to a covered work, for which you have or can give appropriate copyright permission.
Notwithstanding any other provision of this License, for material you add to a covered work, you may (if authorized by the copyright holders of that material) supplement the terms of this License with terms:
a) Disclaiming warranty or limiting liability differently from the terms of sections 15 and 16 of this License; or
b) Requiring preservation of specified reasonable legal notices or author attributions in that material or in the Appropriate Legal Notices displayed by works containing it; or
c) Prohibiting misrepresentation of the origin of that material, or requiring that modified versions of such material be marked in reasonable ways as different from the original version; or
d) Limiting the use for publicity purposes of names of licensors or authors of the material; or
e) Declining to grant rights under trademark law for use of some trade names, trademarks, or service marks; or
f) Requiring indemnification of licensors and authors of that material by anyone who conveys the material (or modified versions of it) with contractual assumptions of liability to the recipient, for any liability that these contractual assumptions directly impose on those licensors and authors.
All other non-permissive additional terms are considered “further restrictions” within the meaning of section 10. If the Program as you received it, or any part of it, contains a notice stating that it is governed by this License along with a term that is a further restriction, you may remove that term. If a license document contains a further restriction but permits relicensing or conveying under this License, you may add to a covered work material governed by the terms of that license document, provided that the further restriction does not survive such relicensing or conveying.
If you add terms to a covered work in accord with this section, you must place, in the relevant source files, a statement of the additional terms that apply to those files, or a notice indicating where to find the applicable terms.
Additional terms, permissive or non-permissive, may be stated in the form of a separately written license, or stated as exceptions; the above requirements apply either way.
8. Termination.
You may not propagate or modify a covered work except as expressly provided under this License. Any attempt otherwise to propagate or modify it is void, and will automatically terminate your rights under this License (including any patent licenses granted under the third paragraph of section 11).
However, if you cease all violation of this License, then your license from a particular copyright holder is reinstated (a) provisionally, unless and until the copyright holder explicitly and finally terminates your license, and (b) permanently, if the copyright holder fails to notify you of the violation by some reasonable means prior to 60 days after the cessation.
Moreover, your license from a particular copyright holder is reinstated permanently if the copyright holder notifies you of the violation by some reasonable means, this is the first time you have received notice of violation of this License (for any work) from that copyright holder, and you cure the violation prior to 30 days after your receipt of the notice.
Termination of your rights under this section does not terminate the licenses of parties who have received copies or rights from you under this License. If your rights have been terminated and not permanently reinstated, you do not qualify to receive new licenses for the same material under section 10.
9. Acceptance Not Required for Having Copies.
You are not required to accept this License in order to receive or run a copy of the Program. Ancillary propagation of a covered work occurring solely as a consequence of using peer-to-peer transmission to receive a copy likewise does not require acceptance. However, nothing other than this License grants you permission to propagate or modify any covered work. These actions infringe copyright if you do not accept this License. Therefore, by modifying or propagating a covered work, you indicate your acceptance of this License to do so.
10. Automatic Licensing of Downstream Recipients.
Each time you convey a covered work, the recipient automatically receives a license from the original licensors, to run, modify and propagate that work, subject to this License. You are not responsible for enforcing compliance by third parties with this License.
An “entity transaction” is a transaction transferring control of an organization, or substantially all assets of one, or subdividing an organization, or merging organizations. If propagation of a covered work results from an entity transaction, each party to that transaction who receives a copy of the work also receives whatever licenses to the work the party's predecessor in interest had or could give under the previous paragraph, plus a right to possession of the Corresponding Source of the work from the predecessor in interest, if the predecessor has it or can get it with reasonable efforts.
You may not impose any further restrictions on the exercise of the rights granted or affirmed under this License. For example, you may not impose a license fee, royalty, or other charge for exercise of rights granted under this License, and you may not initiate litigation (including a cross-claim or counterclaim in a lawsuit) alleging that any patent claim is infringed by making, using, selling, offering for sale, or importing the Program or any portion of it.
11. Patents.
A “contributor” is a copyright holder who authorizes use under this License of the Program or a work on which the Program is based. The work thus licensed is called the contributor's “contributor version”.
A contributor's “essential patent claims” are all patent claims owned or controlled by the contributor, whether already acquired or hereafter acquired, that would be infringed by some manner, permitted by this License, of making, using, or selling its contributor version, but do not include claims that would be infringed only as a consequence of further modification of the contributor version. For purposes of this definition, “control” includes the right to grant patent sublicenses in a manner consistent with the requirements of this License.
Each contributor grants you a non-exclusive, worldwide, royalty-free patent license under the contributor's essential patent claims, to make, use, sell, offer for sale, import and otherwise run, modify and propagate the contents of its contributor version.
In the following three paragraphs, a “patent license” is any express agreement or commitment, however denominated, not to enforce a patent (such as an express permission to practice a patent or covenant not to sue for patent infringement). To “grant” such a patent license to a party means to make such an agreement or commitment not to enforce a patent against the party.
If you convey a covered work, knowingly relying on a patent license, and the Corresponding Source of the work is not available for anyone to copy, free of charge and under the terms of this License, through a publicly available network server or other readily accessible means, then you must either (1) cause the Corresponding Source to be so available, or (2) arrange to deprive yourself of the benefit of the patent license for this particular work, or (3) arrange, in a manner consistent with the requirements of this License, to extend the patent license to downstream recipients. “Knowingly relying” means you have actual knowledge that, but for the patent license, your conveying the covered work in a country, or your recipient's use of the covered work in a country, would infringe one or more identifiable patents in that country that you have reason to believe are valid.
If, pursuant to or in connection with a single transaction or arrangement, you convey, or propagate by procuring conveyance of, a covered work, and grant a patent license to some of the parties receiving the covered work authorizing them to use, propagate, modify or convey a specific copy of the covered work, then the patent license you grant is automatically extended to all recipients of the covered work and works based on it.
A patent license is “discriminatory” if it does not include within the scope of its coverage, prohibits the exercise of, or is conditioned on the non-exercise of one or more of the rights that are specifically granted under this License. You may not convey a covered work if you are a party to an arrangement with a third party that is in the business of distributing software, under which you make payment to the third party based on the extent of your activity of conveying the work, and under which the third party grants, to any of the parties who would receive the covered work from you, a discriminatory patent license (a) in connection with copies of the covered work conveyed by you (or copies made from those copies), or (b) primarily for and in connection with specific products or compilations that contain the covered work, unless you entered into that arrangement, or that patent license was granted, prior to 28 March 2007.
Nothing in this License shall be construed as excluding or limiting any implied license or other defenses to infringement that may otherwise be available to you under applicable patent law.
12. No Surrender of Others' Freedom.
If conditions are imposed on you (whether by court order, agreement or otherwise) that contradict the conditions of this License, they do not excuse you from the conditions of this License. If you cannot convey a covered work so as to satisfy simultaneously your obligations under this License and any other pertinent obligations, then as a consequence you may not convey it at all. For example, if you agree to terms that obligate you to collect a royalty for further conveying from those to whom you convey the Program, the only way you could satisfy both those terms and this License would be to refrain entirely from conveying the Program.
13. Use with the GNU Affero General Public License.
Notwithstanding any other provision of this License, you have permission to link or combine any covered work with a work licensed under version 3 of the GNU Affero General Public License into a single combined work, and to convey the resulting work. The terms of this License will continue to apply to the part which is the covered work, but the special requirements of the GNU Affero General Public License, section 13, concerning interaction through a network will apply to the combination as such.
14. Revised Versions of this License.
The Free Software Foundation may publish revised and/or new versions of the GNU General Public License from time to time. Such new versions will be similar in spirit to the present version, but may differ in detail to address new problems or concerns.
Each version is given a distinguishing version number. If the Program specifies that a certain numbered version of the GNU General Public License “or any later version” applies to it, you have the option of following the terms and conditions either of that numbered version or of any later version published by the Free Software Foundation. If the Program does not specify a version number of the GNU General Public License, you may choose any version ever published by the Free Software Foundation.
If the Program specifies that a proxy can decide which future versions of the GNU General Public License can be used, that proxy's public statement of acceptance of a version permanently authorizes you to choose that version for the Program.
Later license versions may give you additional or different permissions. However, no additional obligations are imposed on any author or copyright holder as a result of your choosing to follow a later version.
15. Disclaimer of Warranty.
THERE IS NO WARRANTY FOR THE PROGRAM, TO THE EXTENT PERMITTED BY APPLICABLE LAW. EXCEPT WHEN OTHERWISE STATED IN WRITING THE COPYRIGHT HOLDERS AND/OR OTHER PARTIES PROVIDE THE PROGRAM “AS IS” WITHOUT WARRANTY OF ANY KIND, EITHER EXPRESSED OR IMPLIED, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE. THE ENTIRE RISK AS TO THE QUALITY AND PERFORMANCE OF THE PROGRAM IS WITH YOU. SHOULD THE PROGRAM PROVE DEFECTIVE, YOU ASSUME THE COST OF ALL NECESSARY SERVICING, REPAIR OR CORRECTION.
16. Limitation of Liability.
IN NO EVENT UNLESS REQUIRED BY APPLICABLE LAW OR AGREED TO IN WRITING WILL ANY COPYRIGHT HOLDER, OR ANY OTHER PARTY WHO MODIFIES AND/OR CONVEYS THE PROGRAM AS PERMITTED ABOVE, BE LIABLE TO YOU FOR DAMAGES, INCLUDING ANY GENERAL, SPECIAL, INCIDENTAL OR CONSEQUENTIAL DAMAGES ARISING OUT OF THE USE OR INABILITY TO USE THE PROGRAM (INCLUDING BUT NOT LIMITED TO LOSS OF DATA OR DATA BEING RENDERED INACCURATE OR LOSSES SUSTAINED BY YOU OR THIRD PARTIES OR A FAILURE OF THE PROGRAM TO OPERATE WITH ANY OTHER PROGRAMS), EVEN IF SUCH HOLDER OR OTHER PARTY HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGES.
17. Interpretation of Sections 15 and 16.
If the disclaimer of warranty and limitation of liability provided above cannot be given local legal effect according to their terms, reviewing courts shall apply local law that most closely approximates an absolute waiver of all civil liability in connection with the Program, unless a warranty or assumption of liability accompanies a copy of the Program in return for a fee.
END OF TERMS AND CONDITIONS
How to Apply These Terms to Your New Programs
If you develop a new program, and you want it to be of the greatest possible use to the public, the best way to achieve this is to make it free software which everyone can redistribute and change under these terms.
To do so, attach the following notices to the program. It is safest to attach them to the start of each source file to most effectively state the exclusion of warranty; and each file should have at least the “copyright” line and a pointer to where the full notice is found.
<one line to give the program's name and a brief idea of what it does.>
Copyright (C) <year> <name of author>
This program is free software: you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation, either version 3 of the License, or (at your option) any later version.
This program is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details.
You should have received a copy of the GNU General Public License along with this program. If not, see <https://www.gnu.org/licenses/>.
Also add information on how to contact you by electronic and paper mail.
If the program does terminal interaction, make it output a short notice like this when it starts in an interactive mode:
<program> Copyright (C) <year> <name of author>
This program comes with ABSOLUTELY NO WARRANTY; for details type `show w'.
This is free software, and you are welcome to redistribute it under certain conditions; type `show c' for details.
The hypothetical commands `show w' and `show c' should show the appropriate parts of the General Public License. Of course, your program's commands might be different; for a GUI interface, you would use an “about box”.
You should also get your employer (if you work as a programmer) or school, if any, to sign a “copyright disclaimer” for the program, if necessary. For more information on this, and how to apply and follow the GNU GPL, see <https://www.gnu.org/licenses/>.
The GNU General Public License does not permit incorporating your program into proprietary programs. If your program is a subroutine library, you may consider it more useful to permit linking proprietary applications with the library. If this is what you want to do, use the GNU Lesser General Public License instead of this License. But first, please read <https://www.gnu.org/philosophy/why-not-lgpl.html>.
+169 -21
View File
@@ -31,9 +31,50 @@ mod intents;
/// Android application entry point, called by android-activity's glue.
#[no_mangle]
fn android_main(app: slint::android::AndroidApp) {
// Before the logger, because the logger needs somewhere to write, and
// before any `log::` call at all, because records emitted before this line
// reach nothing.
// **The first statement in the process, and it has to be.** Everything
// between here and `install` returning runs with no logger installed at
// all: asking the activity for its external directory, `create_dir_all`
// and an `open` on a FUSE-backed volume the system may still be mounting.
// A failure or a stall in any of it is invisible on every surface there
// is — no file yet, and nothing in logcat either — which is precisely the
// kind of launch logcat exists to debug.
//
// `AndroidLogger` rather than `init_once`, so logcat can be *teed* rather
// than replaced: `init_once` installs itself as the global logger and
// there is only one of those. Everything that reached logcat before the
// file existed still reaches it, at the same level and under the same tag;
// the file is strictly additional.
let console = android_logger::AndroidLogger::new(
android_logger::Config::default()
.with_max_level(log::LevelFilter::Info)
.with_tag("DarkRoom"),
);
// Handed to the logger directly, and **not** written as `log::info!`,
// which here would compile and emit nothing: the facade's maximum level is
// `Off` until `diagnostics::install` sets it, and the macro tests that
// before it reaches any logger at all. This call skips the facade and
// reaches `__android_log_write` with nothing in between.
//
// That independence is the second reason for it. When the log is silent,
// this line is what says which half is at fault: present here and absent
// below means the `log` wiring, absent in both means liblog is not
// delivering this process's records — a question about the device, which
// no amount of reading this file can answer.
log::Log::log(
&console,
&log::Record::builder()
.level(log::Level::Info)
.target(module_path!())
.module_path(Some(module_path!()))
.args(format_args!(
"DarkRoom v{} starting; logcat only until the log file opens",
env!("CARGO_PKG_VERSION")
))
.build(),
);
// Before the file logger, because it needs somewhere to write.
//
// **The external directory, not the internal one, and the difference is
// the entire point of the file.** Both are app-private and both survive
@@ -53,16 +94,6 @@ fn android_main(app: slint::android::AndroidApp) {
dr_plat::set_state_dir(dir);
}
// `AndroidLogger` rather than `init_once`, so logcat can be *teed* rather
// than replaced: `init_once` installs itself as the global logger and
// there is only one of those. Everything that reached logcat before this
// change still reaches it, at the same level and under the same tag; the
// file is strictly additional.
let console = android_logger::AndroidLogger::new(
android_logger::Config::default()
.with_max_level(log::LevelFilter::Info)
.with_tag("DarkRoom"),
);
let logging = dr_plat::diagnostics::install(Box::new(console), log::LevelFilter::Info);
// Panics go to stderr, and Android discards stderr. Without this hook a
@@ -121,8 +152,10 @@ fn android_main(app: slint::android::AndroidApp) {
None => log::error!("no internal data path; settings will not persist"),
}
// After the data dir and before anything asks whether a model is present.
install_bundled_models(&app);
// After the data dir, because it writes beside the catalog. **Not** before
// the first frame any more — it starts a worker and returns; see the
// function for what it used to cost the launch.
install_bundled_models(app.clone());
// Before `init_with_event_listener`, which takes `app` by value and is the
// last moment anything can ask the activity a question. Not an ordering
@@ -225,10 +258,58 @@ fn android_main(app: slint::android::AndroidApp) {
/// whether or not the tab is opened. Assets are also *stored* rather than
/// deflated in the APK (see `assemble-apk.sh`), so unpacking is a copy rather
/// than an inflate.
///
/// # Why it returns before it has done anything
///
/// **`android_main` runs with the input channel unserviced.** Nothing drains
/// it until Slint reaches `poll_events`, and Slint does not reach `poll_events`
/// until `dr_ui::run` calls `window.run()`, which is the last line of it. So
/// every millisecond spent between the top of `android_main` and that line is a
/// millisecond in which Android's input dispatcher gets no answer, and five
/// thousand of them is an ANR by definition — the system puts "DarkRoom isn't
/// responding" over a window that has never painted, and offers to kill it.
///
/// This copied **41 MB** on the first launch after an install: 24.9 MB of scene
/// model, 13.6 MB of embedder, 2.5 MB of detector, each read whole out of the
/// APK and written to `/data`. v0.10.0 added the scene model, which was 60% of
/// that total; v0.10.0 is the release the ANR appeared in, and the 8,010 minor
/// faults in its report are what 41 MB of freshly touched pages looks like.
/// The two further detectors the settings page offers since have made it
/// 61 MB, which is the same argument with a larger number.
///
/// So it runs on a worker (NFR-ARCH-1: nothing blocking on the UI executor) and
/// this function returns as soon as the thread is running. Nothing on the
/// launch path waits for it, and no other startup step needs its result.
///
/// # The window in which a model looks absent, and why that is honest enough
///
/// Until the copy finishes, `library::face_models` and `library::scene_model`
/// answer `is_file()` about files that are not written yet, so both report
/// their feature unavailable — the same answer they give a build carrying no
/// weights at all, which is the ordinary case this whole path was written
/// around. It is briefly pessimistic rather than wrong, it lasts about as long
/// as it takes to read one screenful of the grid, and the temporary name
/// [`unpack_bundled_models`] writes under is what stops it being worse than
/// pessimistic: a lookup never sees a half-written file, only an absent one.
#[cfg(target_os = "android")]
fn install_bundled_models(app: &slint::android::AndroidApp) {
fn install_bundled_models(app: slint::android::AndroidApp) {
// Detached rather than joined: there is no later moment on the launch path
// that wants the answer, and a handle nobody joins is a handle nobody can
// forget to. `AndroidApp` is documented `Send` and `Sync` and is an `Arc`
// internally, so the clone costs a refcount; `asset_manager` is asked for
// on the worker because `AAssetManager` is thread-safe by contract and
// reading the pointer takes only the app's read lock, which `poll_events`
// also only ever holds shared.
std::thread::spawn(move || unpack_bundled_models(&app));
}
/// The copy itself, on the worker [`install_bundled_models`] starts.
#[cfg(target_os = "android")]
fn unpack_bundled_models(app: &slint::android::AndroidApp) {
use std::io::Read;
let started = std::time::Instant::now();
// The face names are the **shape-fixed** exports, matching what
// `library::face_models` looks for: tract cannot parse either InsightFace
// graph with its dynamic input dimension, so what ships here has already
@@ -238,25 +319,59 @@ fn install_bundled_models(app: &slint::android::AndroidApp) {
// decodes to 150 anonymous channels — `library::scene_model` wants the
// vocabulary and the category descriptor beside it, and requires all three
// before it reports the tab available.
const BUNDLED: [(&std::ffi::CStr, &str); 5] = [
//
// Three detectors, because which one runs is a setting
// (`FaceDetector`, docs/faces.md §12.3) and a tablet has no other way to
// obtain the one it was not shipped with. Twenty megabytes of APK for
// the choice; the embedder is the same for all three.
//
// Then the three eye-state models (docs/faces.md §17): landmarks, open
// or closed, sunglasses. The app indexes without them; with them the
// eyes-open filter has something to read, and a tablet has no other way
// to get them either.
//
// The int8 forms beside the three detectors are what the Hexagon runs
// (docs/inference.md §5); the engine loads the sibling when the probe
// chose that rung and ignores it otherwise.
const BUNDLED: [(&std::ffi::CStr, &str); 14] = [
(c"models/scrfd_500m_640.onnx", "scrfd_500m_640.onnx"),
(
c"models/scrfd_500m_640.int8.onnx",
"scrfd_500m_640.int8.onnx",
),
(c"models/scrfd_2.5g_640.onnx", "scrfd_2.5g_640.onnx"),
(
c"models/scrfd_2.5g_640.int8.onnx",
"scrfd_2.5g_640.int8.onnx",
),
(c"models/scrfd_10g_640.onnx", "scrfd_10g_640.onnx"),
(c"models/scrfd_10g_640.int8.onnx", "scrfd_10g_640.int8.onnx"),
(c"models/arcface_mbf_b1.onnx", "arcface_mbf_b1.onnx"),
(c"models/2d106det_b1.onnx", "2d106det_b1.onnx"),
(c"models/ocec_s_b1.onnx", "ocec_s_b1.onnx"),
(c"models/sgc_l_48_b1.onnx", "sgc_l_48_b1.onnx"),
(c"models/yolo26s-sem-ade20k.onnx", "yolo26s-sem-ade20k.onnx"),
(
c"models/yolo26s-sem-ade20k.classes.json",
"yolo26s-sem-ade20k.classes.json",
),
(c"models/categories.txt", "categories.txt"),
// The panorama border filler (FR-MRG-4); MIT, 28 MB.
(c"models/migan-512.onnx", "migan-512.onnx"),
];
let dir = dr_ui::shared_face_models_dir();
let assets = app.asset_manager();
let mut copied = 0u64;
for (asset_path, name) in BUNDLED {
let dest = dir.join(name);
// Already unpacked. Not re-read on every launch: this is 15 MB through
// a decompressor on the startup path, and the file does not change
// without the APK changing, at which point the install wiped it anyway.
// Already unpacked. Not re-read on every launch: this is 73 MB of
// copying across the ten entries, and the file does not change without
// the APK changing, at which point the install wiped it anyway. It
// matters more now than it did — a launch that skips every entry here
// costs nothing at all, which is what makes the second launch after an
// install cheap even though the first one is not.
if dest.is_file() {
continue;
}
@@ -281,13 +396,46 @@ fn install_bundled_models(app: &slint::android::AndroidApp) {
// than a missing one.
let part = dir.join(format!("{name}.part"));
match std::fs::write(&part, &bytes).and_then(|()| std::fs::rename(&part, &dest)) {
Ok(()) => log::info!("installed bundled {name} ({} bytes)", bytes.len()),
Ok(()) => {
copied += bytes.len() as u64;
log::info!("installed bundled {name} ({} bytes)", bytes.len());
}
Err(e) => {
log::error!("cannot install {name}: {e}");
let _ = std::fs::remove_file(&part);
}
}
}
// The figure this whole function is about. Said even when it is zero, so a
// launch that ANRs anyway can be told apart from one that spent its six
// seconds here — on a second launch there is nothing left to copy and the
// line reads `0 bytes`.
log::info!(
"bundled models ready: {copied} bytes copied in {} ms",
started.elapsed().as_millis()
);
// Now, and not at launch: the probe fingerprints the model files, and
// on a first launch they were not on disk until this line. The runtime
// is in the APK's native library directory beside `libdarkroom.so`,
// which is also where Qualcomm's DSP loader has to be pointed for the
// Hexagon skel (docs/inference.md §3, §8).
dr_ui::inference::init(native_library_dir().into_iter().collect());
}
/// The directory the system unpacked this APK's native libraries into.
///
/// Read from where the loader put *this* library rather than asked of the
/// activity: `android-activity` does not expose `nativeLibraryDir`, and the
/// answer is in `/proc/self/maps` for free.
#[cfg(target_os = "android")]
fn native_library_dir() -> Option<std::path::PathBuf> {
let maps = std::fs::read_to_string("/proc/self/maps").ok()?;
maps.lines()
.filter_map(|l| l.split_whitespace().nth(5))
.find(|p| p.ends_with("/libdarkroom.so"))
.and_then(|p| std::path::Path::new(p).parent().map(Into::into))
}
/// TRACES: FR-PLAT-AND-6
+8
View File
@@ -15,5 +15,13 @@ anyhow.workspace = true
env_logger.workspace = true
log.workspace = true
# The Windows resource block — icon and version — compiled in by build.rs.
# Unconditional rather than under `[target.'cfg(windows)']`, because a cfg on
# a build-dependency is evaluated against the *host* — the machine running
# the build script — and this is built for Windows from Linux. The script
# itself returns before touching the crate on every other target.
[build-dependencies]
winresource = "0.1"
[features]
default = []
+58
View File
@@ -0,0 +1,58 @@
//! TRACES: FR-PLAT-WIN-2
//! The Windows resource block: icon and version, compiled into the executable.
//!
//! Windows takes an application's icon and its "Details" tab from a resource
//! inside the `.exe`, not from a `.desktop` file, so without this the installed
//! program shows the generic executable icon in Explorer, the Start Menu and
//! the taskbar, and reports no version. Nothing here runs for any other
//! target: the whole body is behind the target-OS check, and the crate that
//! does the work is a build-dependency only.
//!
//! The icon is the same PNG every other platform uses, wrapped into an `.ico`
//! in `OUT_DIR` rather than committed: an ICO entry may *be* a PNG (Vista and
//! later read them directly), so the wrapper is a 22-byte header and the
//! file's bytes, and a generated binary stays out of the tree.
use std::io::Write as _;
use std::path::PathBuf;
fn main() {
println!("cargo:rerun-if-changed=build.rs");
if std::env::var("CARGO_CFG_TARGET_OS").as_deref() != Ok("windows") {
return;
}
let png = PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../../ui/dr-ui/ui/app-icon.png");
println!("cargo:rerun-if-changed={}", png.display());
let bytes = std::fs::read(&png).expect("read app-icon.png");
let ico = PathBuf::from(std::env::var("OUT_DIR").unwrap()).join("darkroom.ico");
write_png_ico(&ico, &bytes, 256).expect("write darkroom.ico");
let mut res = winresource::WindowsResource::new();
res.set_icon(ico.to_str().unwrap());
res.set("ProductName", "DarkRoom");
res.set("FileDescription", "DarkRoom");
res.set("LegalCopyright", "GPL-3.0-or-later");
// Cross-compiling: `winresource` looks for a `windres` for the target and
// the Windows image names it explicitly, for the same reason the Android
// image names its linkers.
if let Ok(windres) = std::env::var("WINDRES") {
res.set_windres_path(&windres);
}
res.compile().expect("compile the Windows resource block");
}
/// One PNG image as an `.ico`. `edge` is the PNG's width and height; 256 is
/// written as 0 per the format.
fn write_png_ico(path: &std::path::Path, png: &[u8], edge: u32) -> std::io::Result<()> {
let mut f = std::fs::File::create(path)?;
let dim = if edge >= 256 { 0u8 } else { edge as u8 };
// ICONDIR: reserved, type 1 (icon), one image.
f.write_all(&[0, 0, 1, 0, 1, 0])?;
// ICONDIRENTRY: width, height, palette 0, reserved, planes 1, bpp 32,
// byte length, offset (6 + 16).
f.write_all(&[dim, dim, 0, 0, 1, 0, 32, 0])?;
f.write_all(&(png.len() as u32).to_le_bytes())?;
f.write_all(&22u32.to_le_bytes())?;
f.write_all(png)
}
+61
View File
@@ -1,12 +1,32 @@
//! DarkRoom desktop entry point.
//!
//! darkroom-desktop <file-or-directory>...
//! darkroom-desktop --version
// TRACES: FR-PLAT-WIN-2
// A GUI-subsystem executable, or Windows opens a console window behind the
// application for the life of the process. Release only: the console is where
// the log goes when there is no file, and a debug build is run from one.
// `--version` still prints under this — stdout is simply not attached when
// launched from Explorer, which is not where anyone asks for a version.
#![cfg_attr(all(windows, not(debug_assertions)), windows_subsystem = "windows")]
use std::path::PathBuf;
use dr_plat::diagnostics::Installed;
fn main() -> anyhow::Result<()> {
// TRACES: FR-PLAT-WIN-3
// Before the logger, the crash hook and everything else: this exists so a
// build made on a machine that cannot run the application — the Linux CI
// producing the Windows binary, checked under Wine — has an exit that
// proves the executable starts without opening a window or touching the
// user's directories (docs/windows.md §6).
if std::env::args().nth(1).as_deref() == Some("--version") {
println!("darkroom-desktop {}", env!("CARGO_PKG_VERSION"));
return Ok(());
}
// Built rather than `init`ed, so the same logger can be handed to the
// diagnostics tee: `env_logger` keeps writing to stderr exactly as before,
// and every record it accepts is also appended to the on-disk log
@@ -41,6 +61,11 @@ fn main() -> anyhow::Result<()> {
eprintln!("usage: darkroom-desktop <file-or-directory>...");
}
// Before the window: the probe runs on its own thread and the first
// frame does not wait for it, but the models a background job asks for
// should already know where the runtime is (docs/inference.md §4).
dr_ui::inference::init(runtime_dirs());
dr_ui::run(paths)?;
// Skip Rust's normal static/thread-local teardown on the way out: a
@@ -50,3 +75,39 @@ fn main() -> anyhow::Result<()> {
// destruction" when the window is closed.
std::process::exit(0);
}
/// Where a desktop package may have put `libonnxruntime`, most specific
/// first. None of these existing is the tract build, which is a complete
/// application and not an error (docs/inference.md §3).
///
/// `DARKROOM_ORT_DIR` is for a developer pointing at a runtime that is not
/// installed — the wheel's `capi` directory, say. Then beside the executable
/// and in the package's private library directory, for a package that
/// bundles its own; then the user's own `runtime/` beside the models, where
/// `tools/fetch-desktop-runtime.sh` puts one; then the Flatpak prefix; then
/// the system library directory, for a distribution that ships ONNX Runtime
/// as a package of its own. The user's copy outranks the system's because
/// the system's is the one most likely to be built without the GPU
/// providers, or against the wrong cuDNN — and a system copy whose providers
/// do not load is not a problem, only a slower app: the probe builds a real
/// session before believing a provider.
fn runtime_dirs() -> Vec<PathBuf> {
let mut dirs = Vec::new();
if let Some(dir) = std::env::var_os("DARKROOM_ORT_DIR") {
dirs.push(PathBuf::from(dir));
}
if let Ok(exe) = std::env::current_exe() {
if let Some(bin) = exe.parent() {
dirs.push(bin.to_path_buf());
dirs.push(bin.join("../lib/darkroom"));
}
}
dirs.push(dr_ui::inference::user_runtime_dir());
#[cfg(target_os = "linux")]
dirs.extend([
PathBuf::from("/app/lib/darkroom"),
PathBuf::from("/usr/lib/darkroom"),
PathBuf::from("/usr/lib"),
]);
dirs
}
+23 -15
View File
@@ -34,6 +34,10 @@ struct Known {
crop_px: f32,
}
/// What the catalog holds per face, decoded: photograph, vector, size,
/// quality.
type Decoded = (u64, Vec<f32>, f32, Option<f32>);
fn main() {
let args: Vec<String> = std::env::args().skip(1).collect();
let Some(path) = args.first() else {
@@ -59,10 +63,10 @@ fn main() {
let model = dr_face::ModelId::new(MODEL_ID.to_string());
let stored = faces::embeddings(conn, MODEL_ID).expect("embeddings");
let mut embedding_of = HashMap::new();
for (id, image, blob, crop_px) in stored {
if let Some(e) = dr_face::Embedding::from_f16_bytes(model.clone(), &blob) {
embedding_of.insert(id, (image.0, e.v.to_vec(), crop_px));
let mut embedding_of: HashMap<faces::FaceId, Decoded> = HashMap::new();
for f in stored {
if let Some(e) = dr_face::Embedding::from_f16_bytes(model.clone(), &f.embedding) {
embedding_of.insert(f.face, (f.image.0, e.v.to_vec(), f.crop_px, f.quality));
}
}
println!("faces with embeddings: {}", embedding_of.len());
@@ -82,7 +86,7 @@ fn main() {
if !f.confirmed {
continue;
}
if let Some((image, embedding, crop_px)) = embedding_of.get(&f.id) {
if let Some((image, embedding, crop_px, _)) = embedding_of.get(&f.id) {
mine.push(Known {
image: *image,
person: p.id,
@@ -288,19 +292,22 @@ fn band(label: &str, v: &[f32]) {
/// The whole library through the real clusterer, for the numbers it would
/// actually write.
fn full_library(
embedding_of: &HashMap<faces::FaceId, (u64, Vec<f32>, f32)>,
embedding_of: &HashMap<faces::FaceId, Decoded>,
confirmed: &HashMap<faces::FaceId, u64>,
cal: &dr_face::Calibration,
) {
let mut candidates: Vec<dr_face::Candidate> = embedding_of
.iter()
.map(|(id, (image, embedding, crop_px))| dr_face::Candidate {
face: id.0,
image: *image,
embedding: embedding.clone(),
crop_px: *crop_px,
confirmed_person: confirmed.get(id).copied(),
})
.map(
|(id, (image, embedding, crop_px, quality))| dr_face::Candidate {
face: id.0,
image: *image,
embedding: embedding.clone(),
crop_px: *crop_px,
quality: *quality,
confirmed_person: confirmed.get(id).copied(),
},
)
.collect();
candidates.sort_by_key(|c| c.face);
@@ -317,11 +324,13 @@ fn full_library(
.collect();
let crop_px: Vec<f32> = candidates.iter().map(|c| c.crop_px).collect();
let images: Vec<u64> = candidates.iter().map(|c| c.image).collect();
let gallery: Vec<bool> = candidates.iter().map(|c| c.in_gallery()).collect();
let view = dr_face::neighbours::Faces {
embeddings: &flat,
dim,
crop_px: &crop_px,
images: &images,
gallery: &gallery,
};
let t = std::time::Instant::now();
@@ -335,8 +344,7 @@ fn full_library(
let agglomerate = t.elapsed().as_secs_f64() - scan;
let t = std::time::Instant::now();
let _ =
dr_face::identity_shares(candidates.len(), &clusters, &evidence, dr_face::TOP_MATCHES);
let _ = dr_face::identity_shares(&gallery, &clusters, &evidence, dr_face::TOP_MATCHES);
println!(
" scan {scan:.2}s ({} evidence pairs) · agglomerate {agglomerate:.2}s · score {:.2}s",
evidence.len(),
+218
View File
@@ -535,6 +535,105 @@ pub fn collections_for_image(
Ok(rows)
}
/// One collection a set of images is filed in, and how much of that set is in
/// it.
///
/// `holding` is what makes a removal honest: with forty photographs selected
/// and three of them in "Iceland", the row has to say "3 of 40" or the user
/// reads it as "this selection is in Iceland" and takes all forty out of a
/// collection thirty-seven of them were never in.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct Membership {
pub id: CollectionId,
pub name: String,
/// How many of the images asked about are members. Never zero — a
/// collection holding none of them is not returned at all.
pub holding: usize,
}
/// Every collection the given images are filed in, with how many of them each
/// holds.
///
/// The read behind "which collections is this selection in, and take it out of
/// one" — the counterpart to [`collections_for_image`], which answers the same
/// question for a single photograph and does not need the counts.
///
/// Smart collections never appear: they have no `collection_members` rows, so
/// there is nothing to remove and offering it would be a button that does
/// nothing. Sorted by name, matching the sidebar.
///
/// # Why this is chunked
///
/// The image list is a *selection*, which a select-all makes as large as the
/// library. SQLite caps the number of bound parameters in one statement, so a
/// single `IN (...)` over every selected id fails outright on exactly the
/// gesture most likely to produce it. The counts are summed across chunks
/// rather than re-queried, so the result is the same as the unchunked query
/// would have given.
pub fn membership_of(
conn: &Connection,
images: &[ImageId],
) -> Result<Vec<Membership>, CatalogError> {
if images.is_empty() {
return Ok(Vec::new());
}
// Well under SQLite's default parameter cap, and large enough that an
// ordinary selection is one round trip.
const CHUNK: usize = 400;
let mut totals: std::collections::HashMap<CollectionId, (String, usize)> =
std::collections::HashMap::new();
for chunk in images.chunks(CHUNK) {
let placeholders = std::iter::repeat_n("?", chunk.len())
.collect::<Vec<_>>()
.join(",");
// The placeholder list is built from the id *count*, never from user
// text — the same construction `deep_count` uses.
let sql = format!(
"SELECT c.id, c.name, count(*)
FROM collection_members m
JOIN collections c ON c.id = m.collection_id
WHERE m.image_id IN ({placeholders}) AND c.deleted = 0
GROUP BY c.id, c.name"
);
let params: Vec<rusqlite::types::Value> = chunk
.iter()
.map(|i| rusqlite::types::Value::Integer(i.0 as i64))
.collect();
let mut stmt = conn.prepare(&sql)?;
let rows = stmt.query_map(rusqlite::params_from_iter(params.iter()), |r| {
Ok((
CollectionId(r.get::<_, i64>(0)? as u64),
r.get::<_, String>(1)?,
r.get::<_, i64>(2)? as usize,
))
})?;
for row in rows {
let (id, name, n) = row?;
let entry = totals.entry(id).or_insert((name, 0));
entry.1 += n;
}
}
let mut out: Vec<Membership> = totals
.into_iter()
.map(|(id, (name, holding))| Membership { id, name, holding })
.collect();
// By name, then by id, so two collections sharing a name have a stable
// order rather than the hash map's.
out.sort_by(|a, b| {
a.name
.to_lowercase()
.cmp(&b.name.to_lowercase())
.then(a.id.0.cmp(&b.id.0))
});
Ok(out)
}
/// What kind of collection `id` is, or `None` if there is no such collection.
///
/// Cheaper than reading the whole [`Collection`] where the caller only needs to
@@ -552,6 +651,44 @@ pub fn kind(conn: &Connection, id: CollectionId) -> Result<Option<CollectionKind
Ok(found.map(CollectionKind::from_i64))
}
/// TRACES: FR-UI-8
/// The device-independent name of a collection, from its local id.
///
/// The pair to [`id_for_uuid`], and the reason both exist: `collections.id` is
/// an autoincrement local to one catalog, so anything that travels between
/// devices — a place, a merge — has to say which collection it means in the
/// only vocabulary they share.
///
/// `None` for a collection that is not there, or has been tombstoned. A caller
/// writing down a scope treats that as "the whole library", which is the
/// harmless direction: the alternative is recording a name nothing can resolve.
pub fn uuid_of(conn: &Connection, id: CollectionId) -> Result<Option<String>, CatalogError> {
Ok(conn
.query_row(
"SELECT uuid FROM collections WHERE id = ?1 AND deleted = 0",
[id.0 as i64],
|r| r.get::<_, String>(0),
)
.optional()?)
}
/// TRACES: FR-UI-8
/// The local id of a collection, from the name every device knows it by.
///
/// `None` where this device has never heard of it, or has deleted it — a place
/// recorded on the tablet inside a collection this machine has not yet merged.
/// The caller falls back to the whole library rather than to an empty grid.
pub fn id_for_uuid(conn: &Connection, uuid: &str) -> Result<Option<CollectionId>, CatalogError> {
Ok(conn
.query_row(
"SELECT id FROM collections WHERE uuid = ?1 AND deleted = 0",
[uuid],
|r| r.get::<_, i64>(0),
)
.optional()?
.map(|id| CollectionId(id as u64)))
}
/// A collection and everything beneath it, including itself.
///
/// Used for cycle checks and for scoping the grid to a parent: selecting a
@@ -1400,4 +1537,85 @@ mod tests {
let d = descendants(c, a).unwrap();
assert!(d.len() <= 2);
}
#[test]
fn membership_says_how_many_of_the_selection_each_collection_holds() {
// The count is the whole point: "3 of 40" is what stops a user taking
// forty photographs out of a collection thirty-seven were never in.
let cat = seeded();
let c = cat.connection();
let iceland = create(c, "Iceland", None, CollectionKind::Manual).unwrap();
let best = create(c, "Best", None, CollectionKind::Manual).unwrap();
add_images(c, iceland, &[img(1), img(2), img(3)]).unwrap();
add_images(c, best, &[img(1)]).unwrap();
let m = membership_of(c, &[img(1), img(2), img(3), img(4)]).unwrap();
assert_eq!(m.len(), 2);
// Sorted by name, so "Best" comes before "Iceland".
assert_eq!(m[0].id, best);
assert_eq!(m[0].holding, 1);
assert_eq!(m[1].id, iceland);
assert_eq!(m[1].holding, 3);
}
#[test]
fn membership_omits_collections_holding_none_of_them() {
// A row offering to remove images that are not there would be a button
// that does nothing, which is worse than an absent one.
let cat = seeded();
let c = cat.connection();
let a = create(c, "A", None, CollectionKind::Manual).unwrap();
create(c, "Empty", None, CollectionKind::Manual).unwrap();
add_images(c, a, &[img(1)]).unwrap();
let m = membership_of(c, &[img(1)]).unwrap();
assert_eq!(m.len(), 1);
assert_eq!(m[0].id, a);
}
#[test]
fn membership_of_nothing_is_nothing() {
let cat = seeded();
assert!(membership_of(cat.connection(), &[]).unwrap().is_empty());
}
#[test]
fn membership_skips_a_deleted_collection() {
// The tombstone survives the delete, and its member rows are dropped —
// but a merge can leave rows behind, and the sheet must not offer a
// collection the sidebar does not draw.
let cat = seeded();
let c = cat.connection();
let a = create(c, "A", None, CollectionKind::Manual).unwrap();
add_images(c, a, &[img(1)]).unwrap();
delete(c, a).unwrap();
assert!(membership_of(c, &[img(1)]).unwrap().is_empty());
}
#[test]
fn membership_sums_across_chunks() {
// The chunking exists for a select-all, which is exactly the gesture
// that would otherwise exceed SQLite's parameter cap. A count that was
// per-chunk rather than summed would under-report on the one selection
// large enough to need it.
let cat = seeded();
let c = cat.connection();
let a = create(c, "A", None, CollectionKind::Manual).unwrap();
// Past the fixture's six, so the member rows have images to point at.
for i in 7..=900i64 {
c.execute(
"INSERT INTO images(id, root_id, source_ref, added_at)
VALUES (?1, 1, ?2, 0)",
rusqlite::params![i, format!("img{i}.CR3")],
)
.unwrap();
}
let ids: Vec<ImageId> = (1..=900).map(img).collect();
add_images(c, a, &ids).unwrap();
let m = membership_of(c, &ids).unwrap();
assert_eq!(m.len(), 1);
assert_eq!(m[0].holding, 900, "counted across every chunk");
}
}
+341 -44
View File
@@ -48,9 +48,10 @@ pub const SHARD_MAX_BYTES: u64 = dr_thumbs::SHARD_MAX_BYTES;
/// Bytes one stored face occupies, near enough to bound a shard by.
///
/// Counted rather than measured: the embedding is fixed at 512 × f16, the
/// landmarks at 5 × 2 × f32, and the rest is a handful of numbers. Measuring
/// landmarks at 5 × 2 × f32 and the dense ones at 106 × 2 × u16, and the
/// rest is a handful of numbers. Measuring
/// the file after each insert would mean a `VACUUM` to get an honest answer.
const BYTES_PER_FACE: u64 = 1024 + 40 + 64;
const BYTES_PER_FACE: u64 = 1024 + 40 + 424 + 64;
/// Bytes a stored crop occupies, near enough to bound a shard by.
///
@@ -76,6 +77,14 @@ pub struct SharedFace {
pub confidence: f32,
pub embedding: Vec<u8>,
pub crop_px: f32,
/// See `faces::DetectedFace::quality`. `None` from a shard written before
/// the number was kept.
pub quality: Option<f32>,
/// See `faces::DetectedFace::eyes`. `None` from a peer without the eye
/// models, or a shard written before they existed.
pub eyes: Option<dr_face::EyeReading>,
/// See `faces::DetectedFace::landmarks_dense`; empty where none.
pub landmarks_dense: Vec<u8>,
/// The face cut out and encoded, or empty where none was kept.
///
/// Travels with the face rather than in the catalog snapshot, which is the
@@ -146,6 +155,29 @@ impl FaceShardStore {
.flatten()
}
/// The pipeline this store holds an image under, among those sharing
/// `model_id`'s embedder — the most recently indexed where a peer has
/// sent more than one.
///
/// What the import asks: not "has anyone run *this* detector over it" but
/// "does anyone hold comparable faces for it". See `faces::embedder_of`.
pub fn held_model(&self, file_id: u64, model_id: &str) -> Option<String> {
self.index
.query_row(
&format!(
"SELECT model_id FROM entries
WHERE file_id = ?1 AND {} = ?2
ORDER BY indexed_at DESC NULLS LAST, model_id",
crate::faces::embedder_sql("model_id")
),
rusqlite::params![file_id as i64, crate::faces::embedder_of(model_id)],
|r| r.get::<_, String>(0),
)
.optional()
.ok()
.flatten()
}
pub fn contains(&self, file_id: u64, model_id: &str) -> bool {
self.index
.query_row(
@@ -224,8 +256,12 @@ impl FaceShardStore {
tx.execute(
"INSERT INTO faces
(file_id, model_id, x, y, w, h, landmarks, confidence,
embedding, crop_px, crop)
VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11)",
embedding, crop_px, crop, quality,
eye_right, eye_right_px, eye_right_sharp,
eye_left, eye_left_px, eye_left_sharp, sunglasses,
landmarks_dense)
VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11, ?12,
?13, ?14, ?15, ?16, ?17, ?18, ?19, ?20)",
rusqlite::params![
f.file_id as i64,
f.model_id,
@@ -238,6 +274,15 @@ impl FaceShardStore {
f.embedding,
f.crop_px as f64,
(!f.crop.is_empty()).then_some(f.crop.as_slice()),
f.quality.map(f64::from),
f.eyes.map(|e| f64::from(e.right.open)),
f.eyes.map(|e| f64::from(e.right.px)),
f.eyes.map(|e| f64::from(e.right.sharpness)),
f.eyes.map(|e| f64::from(e.left.open)),
f.eyes.map(|e| f64::from(e.left.px)),
f.eyes.map(|e| f64::from(e.left.sharpness)),
f.eyes.map(|e| f64::from(e.sunglasses)),
(!f.landmarks_dense.is_empty()).then_some(f.landmarks_dense.as_slice()),
],
)?;
}
@@ -442,9 +487,16 @@ impl FaceShardStore {
}
let mut fq = src.prepare(&format!(
"SELECT f.file_id, f.model_id, f.x, f.y, f.w, f.h, f.landmarks,
f.confidence, f.embedding, f.crop_px, {}
f.confidence, f.embedding, f.crop_px, {}, {}, {}, {}
FROM faces f WHERE f.file_id = ?1 AND f.model_id = ?2",
crop_column(&src)
column_or_null(&src, "crop"),
column_or_null(&src, "quality"),
crate::schema::EYE_COLUMNS
.iter()
.map(|c| column_or_null(&src, c))
.collect::<Vec<_>>()
.join(", "),
column_or_null(&src, "landmarks_dense"),
))?;
let faces: Vec<SharedFace> = fq
.query_map(rusqlite::params![file_id, &model_id], read_shared_face)?
@@ -484,7 +536,9 @@ impl FaceShardStore {
let Some(edge) = edge else { return Ok(None) };
let mut q = conn.prepare(
"SELECT file_id, model_id, x, y, w, h, landmarks, confidence, embedding, crop_px, crop
"SELECT file_id, model_id, x, y, w, h, landmarks, confidence, embedding, crop_px,
crop, quality, eye_right, eye_right_px, eye_right_sharp,
eye_left, eye_left_px, eye_left_sharp, sunglasses, landmarks_dense
FROM faces WHERE file_id = ?1 AND model_id = ?2",
)?;
let faces: Vec<SharedFace> = q
@@ -541,6 +595,15 @@ fn upgrade_shard(conn: &Connection) -> Result<(), CatalogError> {
for (table, column, decl) in [
("faces", "crop", "BLOB"),
("indexed", "indexed_at", "INTEGER"),
("faces", "quality", "REAL"),
("faces", "eye_right", "REAL"),
("faces", "eye_right_px", "REAL"),
("faces", "eye_right_sharp", "REAL"),
("faces", "eye_left", "REAL"),
("faces", "eye_left_px", "REAL"),
("faces", "eye_left_sharp", "REAL"),
("faces", "sunglasses", "REAL"),
("faces", "landmarks_dense", "BLOB"),
] {
if !has_column(conn, table, column)? {
conn.execute_batch(&format!("ALTER TABLE {table} ADD COLUMN {column} {decl}"))?;
@@ -555,16 +618,19 @@ fn has_column(conn: &Connection, table: &str, column: &str) -> Result<bool, Cata
Ok(stmt.exists(rusqlite::params![table, column])?)
}
/// `f.crop`, or a `NULL` standing in for it.
/// `f.<column>`, or a `NULL` standing in for it.
///
/// A shard downloaded from a peer is opened **read-only** and cannot be
/// upgraded, so one written before crops existed has to be read as it is rather
/// than repaired. Selecting a literal keeps the column count the same, which is
/// what lets [`read_shared_face`] stay a single function.
fn crop_column(conn: &Connection) -> &'static str {
match has_column(conn, "faces", "crop") {
Ok(true) => "f.crop",
_ => "NULL",
/// upgraded, so one written before a column existed has to be read as it is
/// rather than repaired. Selecting a literal keeps the column count the same,
/// which is what lets [`read_shared_face`] stay a single function.
///
/// `column` is one of this module's own names, never anything read from
/// outside, which is what makes formatting it into SQL acceptable.
fn column_or_null(conn: &Connection, column: &str) -> String {
match has_column(conn, "faces", column) {
Ok(true) => format!("f.{column}"),
_ => "NULL".to_string(),
}
}
@@ -598,16 +664,21 @@ pub fn export_to_shards_reporting(
model_id: &str,
progress: &mut dyn FnMut(usize, usize),
) -> Result<usize, CatalogError> {
let mut q = conn.prepare(
"SELECT r.file_id, fi.image_id, fi.source_edge, fi.indexed_at
// Every pipeline sharing this one's embedder, each image under the id
// that actually indexed it. A device that switched detectors still holds
// most of its library under the previous id, and those faces are exactly
// as comparable — and as wanted by a peer — as the new ones.
let mut q = conn.prepare(&format!(
"SELECT r.file_id, fi.image_id, fi.source_edge, fi.indexed_at, fi.model_id
FROM face_index fi
JOIN remote r ON r.image_id = fi.image_id
WHERE fi.model_id = ?1
WHERE {} = ?1
ORDER BY fi.image_id",
)?;
let rows: Vec<(i64, i64, i64, i64)> = q
.query_map([model_id], |r| {
Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?))
crate::faces::embedder_sql("fi.model_id")
))?;
let rows: Vec<(i64, i64, i64, i64, String)> = q
.query_map([crate::faces::embedder_of(model_id)], |r| {
Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?, r.get(4)?))
})?
.collect::<Result<_, _>>()?;
@@ -617,7 +688,8 @@ pub fn export_to_shards_reporting(
let total = rows.len();
let mut exported = 0;
for (seen, (file_id, image_id, edge, indexed_at)) in rows.into_iter().enumerate() {
for (seen, (file_id, image_id, edge, indexed_at, model_id)) in rows.into_iter().enumerate() {
let model_id = model_id.as_str();
if seen.is_multiple_of(REPORT_EVERY) {
progress(seen, total);
}
@@ -637,7 +709,9 @@ pub fn export_to_shards_reporting(
continue;
}
let mut fq = conn.prepare(
"SELECT x, y, w, h, landmarks, detector_confidence, embedding, crop_px, crop
"SELECT x, y, w, h, landmarks, detector_confidence, embedding, crop_px, crop,
quality, eye_right, eye_right_px, eye_right_sharp,
eye_left, eye_left_px, eye_left_sharp, sunglasses, landmarks_dense
FROM faces WHERE image_id = ?1 AND model_id = ?2",
)?;
let faces: Vec<SharedFace> = fq
@@ -654,6 +728,9 @@ pub fn export_to_shards_reporting(
embedding: r.get(6)?,
crop_px: r.get::<_, f64>(7)? as f32,
crop: r.get::<_, Option<Vec<u8>>>(8)?.unwrap_or_default(),
quality: r.get::<_, Option<f64>>(9)?.map(|q| q as f32),
eyes: crate::faces::read_eyes(r, 10)?,
landmarks_dense: r.get::<_, Option<Vec<u8>>>(17)?.unwrap_or_default(),
})
})?
.collect::<Result<_, _>>()?;
@@ -677,10 +754,20 @@ pub fn export_to_shards_reporting(
/// adopted rather than re-detected, which is the difference between a new
/// device being useful in a minute and in two hours.
///
/// Skips any image this device has already indexed itself. Local work is not
/// second-guessed by a peer's — the two should agree, since the same model over
/// the same proxy is deterministic, but where they do not, the copy this device
/// computed is the one it can vouch for.
/// Skips any image this device has already indexed itself under this
/// pipeline or any sharing its embedder — unless the peer ran a detector that
/// outranks the one that indexed it here. Local work is not second-guessed
/// by a peer's equal: the two should agree, since the same model over the
/// same proxy is deterministic, and where they do not, the copy this device
/// computed is the one it can vouch for. A peer's *stronger* pass is another
/// matter: it is the re-detection this device's own sweep would queue
/// (`FaceDetector::supersedes`), already done, and taking it is what spares
/// a tablet the fetch. Names survive the replacement by box overlap and
/// embedding, as they do a local re-detection (`faces::record_detections`).
///
/// A peer's faces are taken under whichever compatible detector found them:
/// a tablet set to the fast detector adopts the desktop's thorough pass
/// rather than re-detecting it worse.
///
/// Returns how many images were adopted.
pub fn import_from_shards(
@@ -688,28 +775,65 @@ pub fn import_from_shards(
store: &FaceShardStore,
model_id: &str,
) -> Result<usize, CatalogError> {
use dr_types::FaceDetector;
// Only images this device actually has. A shard covers the whole account,
// and a device holding a subset of the library should take only its own
// part rather than accumulating faces for photographs it cannot show.
let mut q = conn.prepare(
"SELECT r.file_id, r.image_id
//
// With the pipeline that indexed each one here, or NULL: the marker is
// what decides whether a peer's copy is a gap filled or an upgrade.
let mut q = conn.prepare(&format!(
"SELECT r.file_id, r.image_id,
(SELECT fi.model_id FROM face_index fi
WHERE fi.image_id = r.image_id AND {} = ?1)
FROM remote r
JOIN images i ON i.id = r.image_id
WHERE i.trashed_at IS NULL
AND NOT EXISTS (
SELECT 1 FROM face_index fi
WHERE fi.image_id = r.image_id AND fi.model_id = ?1
)",
)?;
let candidates: Vec<(i64, i64)> = q
.query_map([model_id], |r| Ok((r.get(0)?, r.get(1)?)))?
WHERE i.trashed_at IS NULL",
crate::faces::embedder_sql("fi.model_id")
))?;
let candidates: Vec<(i64, i64, Option<String>)> = q
.query_map([crate::faces::embedder_of(model_id)], |r| {
Ok((r.get(0)?, r.get(1)?, r.get(2)?))
})?
.collect::<Result<_, _>>()?;
let mut adopted = 0;
for (file_id, image_id) in candidates {
let Some((faces, edge)) = store.get_image(file_id as u64, model_id)? else {
for (file_id, image_id, local) in candidates {
let Some(held) = store.held_model(file_id as u64, model_id) else {
continue;
};
if let Some(local) = local {
// An unknown detector on either side cannot be ranked, and an
// unranked peer is treated as an equal: kept out.
let upgrade = match (
FaceDetector::for_model_id(&held),
FaceDetector::for_model_id(&local),
) {
(Some(theirs), Some(ours)) => theirs.outranks(ours),
_ => false,
};
if !upgrade {
continue;
}
}
let Some((faces, edge)) = store.get_image(file_id as u64, &held)? else {
continue;
};
// A peer that embedded before the quality was kept has done work this
// device cannot finish: the number exists only at embedding time, and
// adopting the faces would write the run marker that keeps them from
// ever being measured (schema V14). Left for this device's own pass —
// or for the peer's, whose re-export replaces these.
//
// A missing *eye* reading is not the same case and is adopted. The
// measuring pass finds those by the NULL, not by the marker, so
// adopting the faces costs the reading nothing (schema V16) — and a
// peer that has no eye models may be the only one that has done the
// detection at all.
if faces.iter().any(|f| f.quality.is_none()) {
continue;
}
let local: Vec<crate::faces::DetectedFace> = faces
.into_iter()
.map(|f| crate::faces::DetectedFace {
@@ -721,6 +845,9 @@ pub fn import_from_shards(
confidence: f.confidence,
embedding: f.embedding,
crop_px: f.crop_px,
quality: f.quality,
eyes: f.eyes,
landmarks_dense: f.landmarks_dense,
model_id: f.model_id,
// A peer that indexed before crops existed sends none, and the
// reader falls back to the proxy exactly as it does for a face
@@ -732,7 +859,7 @@ pub fn import_from_shards(
crate::faces::record_detections(
conn,
dr_types::ImageId(image_id as u64),
model_id,
&held,
edge,
&local,
)?;
@@ -766,6 +893,9 @@ fn read_shared_face(r: &rusqlite::Row<'_>) -> rusqlite::Result<SharedFace> {
embedding: r.get(8)?,
crop_px: r.get::<_, f64>(9)? as f32,
crop: r.get::<_, Option<Vec<u8>>>(10)?.unwrap_or_default(),
quality: r.get::<_, Option<f64>>(11)?.map(|q| q as f32),
eyes: crate::faces::read_eyes(r, 12)?,
landmarks_dense: r.get::<_, Option<Vec<u8>>>(19)?.unwrap_or_default(),
})
}
@@ -849,7 +979,24 @@ CREATE TABLE IF NOT EXISTS faces (
crop_px REAL NOT NULL,
-- The face, cut out. NULL where the face was found before crops were kept,
-- or adopted from a peer that did not have one.
crop BLOB
crop BLOB,
-- Length of the raw embedding (`faces::DetectedFace::quality`). NULL from
-- a build that did not keep it, and a face the receiving device will not
-- adopt -- see `import_from_shards`.
quality REAL,
-- The eye reading (`faces::DetectedFace::eyes`), all seven or none. NULL
-- from a peer without the eye models; adopted anyway, and read by the
-- receiving device's own measuring pass if it has them.
eye_right REAL,
eye_right_px REAL,
eye_right_sharp REAL,
eye_left REAL,
eye_left_px REAL,
eye_left_sharp REAL,
sunglasses REAL,
-- The dense landmarks behind the reading (`faces::DetectedFace::
-- landmarks_dense`), 424 bytes packed; NULL where none.
landmarks_dense BLOB
);
CREATE INDEX IF NOT EXISTS faces_file ON faces(file_id, model_id);
@@ -908,6 +1055,9 @@ mod tests {
confidence: 0.87,
embedding: vec![seed; 1024],
crop_px: 180.0,
quality: Some(17.5),
eyes: None,
landmarks_dense: Vec::new(),
crop: vec![seed; 64],
}
}
@@ -925,6 +1075,7 @@ mod tests {
assert_eq!(edge, 1024);
assert_eq!(faces[0].embedding.len(), 1024);
assert!((faces[0].crop_px - 180.0).abs() < 1e-3);
assert_eq!(faces[0].quality, Some(17.5));
}
/// The case the run marker exists for, carried across the wire: an image
@@ -1115,6 +1266,9 @@ mod catalog_round_trip {
confidence: 0.9,
embedding: vec![seed; 1024],
crop_px: 180.0,
quality: Some(20.0),
eyes: None,
landmarks_dense: Vec::new(),
model_id: "w600k_mbf".into(),
crop: vec![seed; 64],
}
@@ -1162,9 +1316,92 @@ mod catalog_round_trip {
let got = faces::for_image(&b, dr_types::ImageId(90)).unwrap();
assert_eq!(got.len(), 1);
assert!((got[0].crop_px - 180.0).abs() < 1e-3);
assert_eq!(got[0].quality, Some(20.0));
assert!((got[0].landmarks[2].0 - 0.15).abs() < 1e-5);
let emb = faces::embeddings(&b, "w600k_mbf").unwrap();
assert!(emb.iter().any(|(_, _, blob, _)| blob[0] == 1));
assert!(emb.iter().any(|e| e.embedding[0] == 1));
}
/// The desktop switched to a stronger detector part-way through the
/// library, so its faces sit under two pipeline ids. A tablet on the
/// original detector must receive *all* of them — each under the id that
/// found it — and not re-detect the thorough half worse.
#[test]
fn every_generation_sharing_an_embedder_travels_and_is_adopted() {
let a = device(&[(1, 5001), (2, 5002)]);
let b = device(&[(90, 5001), (91, 5002)]);
faces::record_detections(&a, dr_types::ImageId(1), "w600k_mbf", 1024, &[detected(1)])
.unwrap();
let mut thorough = detected(2);
thorough.model_id = "scrfd_10g+w600k_mbf".into();
faces::record_detections(
&a,
dr_types::ImageId(2),
"scrfd_10g+w600k_mbf",
1024,
&[thorough],
)
.unwrap();
let mut store_a = FaceShardStore::open(&tempdir("a")).unwrap();
assert_eq!(
export_to_shards(&a, &mut store_a, "scrfd_10g+w600k_mbf").unwrap(),
2,
"the export left the earlier detector's images behind"
);
let mut store_b = FaceShardStore::open(&tempdir("b")).unwrap();
store_b.merge_shard(&store_a.shard_path(0)).unwrap();
assert_eq!(import_from_shards(&b, &store_b, "w600k_mbf").unwrap(), 2);
assert_eq!(faces::coverage(&b, "w600k_mbf").unwrap().outstanding(), 0);
let old = faces::for_image(&b, dr_types::ImageId(90)).unwrap();
let new = faces::for_image(&b, dr_types::ImageId(91)).unwrap();
assert_eq!(old[0].model_id, "w600k_mbf");
assert_eq!(
new[0].model_id, "scrfd_10g+w600k_mbf",
"adopted under the wrong id"
);
}
/// A face a peer embedded without measuring it is work this device
/// cannot finish, and adopting it would write the marker that stops it
/// ever being measured. The image stays outstanding instead.
#[test]
fn a_peers_unmeasured_faces_are_left_for_this_device_to_index() {
let b = device(&[(90, 5001), (91, 5002)]);
let mut store = FaceShardStore::open(&tempdir("unmeasured")).unwrap();
let shared = |file_id: u64, quality: Option<f32>| SharedFace {
file_id,
model_id: "w600k_mbf".into(),
x: 0.1,
y: 0.2,
w: 0.15,
h: 0.2,
landmarks: vec![1; 40],
confidence: 0.87,
embedding: vec![1; 1024],
crop_px: 180.0,
quality,
eyes: None,
landmarks_dense: Vec::new(),
crop: Vec::new(),
};
store
.put_image(5001, "w600k_mbf", 2560, &[shared(5001, None)])
.unwrap();
store
.put_image(5002, "w600k_mbf", 2560, &[shared(5002, Some(19.0))])
.unwrap();
assert_eq!(import_from_shards(&b, &store, "w600k_mbf").unwrap(), 1);
let cov = faces::coverage(&b, "w600k_mbf").unwrap();
assert_eq!(cov.indexed, 1);
assert_eq!(cov.outstanding(), 1, "the unmeasured image was adopted");
assert!(faces::for_image(&b, dr_types::ImageId(90))
.unwrap()
.is_empty());
}
#[test]
@@ -1187,7 +1424,60 @@ mod catalog_round_trip {
"a peer's copy replaced work this device had already done"
);
let emb = faces::embeddings(&b, "w600k_mbf").unwrap();
assert_eq!(emb[0].2[0], 9, "B's own embedding was overwritten");
assert_eq!(emb[0].embedding[0], 9, "B's own embedding was overwritten");
}
/// A peer's stronger detector is the re-detection this device would
/// otherwise queue for itself. Taking it saves the fetch; the name the
/// user confirmed here rides across on the box, as it would locally.
#[test]
fn a_peers_stronger_pass_replaces_a_weaker_local_one_and_keeps_the_name() {
let a = device(&[(1, 5001)]);
let b = device(&[(50, 5001)]);
let ids =
faces::record_detections(&b, dr_types::ImageId(50), "w600k_mbf", 1024, &[detected(9)])
.unwrap();
let anna = faces::create_person(&b, "Anna").unwrap();
faces::confirm(&b, ids[0], anna).unwrap();
let mut thorough = detected(7);
thorough.model_id = "scrfd_10g+w600k_mbf".into();
let mut second = detected(8);
second.model_id = "scrfd_10g+w600k_mbf".into();
second.x = 0.6;
faces::record_detections(
&a,
dr_types::ImageId(1),
"scrfd_10g+w600k_mbf",
1024,
&[thorough, second],
)
.unwrap();
let mut store = FaceShardStore::open(&tempdir("upgrade")).unwrap();
export_to_shards(&a, &mut store, "scrfd_10g+w600k_mbf").unwrap();
assert_eq!(import_from_shards(&b, &store, "w600k_mbf").unwrap(), 1);
let got = faces::for_image(&b, dr_types::ImageId(50)).unwrap();
assert_eq!(got.len(), 2, "the stronger pass was not adopted");
let named = got
.iter()
.find(|f| f.person == Some(anna))
.expect("the name was lost");
assert!(named.confirmed);
assert_eq!(named.model_id, "scrfd_10g+w600k_mbf");
// And never downwards: A on the fast detector keeps B's thorough faces.
let mut store_b = FaceShardStore::open(&tempdir("downgrade")).unwrap();
faces::record_detections(&b, dr_types::ImageId(50), "w600k_mbf", 1024, &[detected(9)])
.unwrap();
export_to_shards(&b, &mut store_b, "w600k_mbf").unwrap();
assert_eq!(
import_from_shards(&a, &store_b, "scrfd_10g+w600k_mbf").unwrap(),
0
);
assert_eq!(faces::for_image(&a, dr_types::ImageId(1)).unwrap().len(), 2);
}
/// A device holding a subset of the library takes only its own part.
@@ -1281,6 +1571,9 @@ mod catalog_round_trip {
confidence: 0.87,
embedding: vec![seed; 1024],
crop_px: 180.0,
quality: None,
eyes: None,
landmarks_dense: Vec::new(),
crop: vec![seed; 64],
}
}
@@ -1430,6 +1723,10 @@ mod catalog_round_trip {
let (faces, _) = store.get_image(77, "w600k_mbf").unwrap().unwrap();
assert_eq!(faces.len(), 1);
assert!(faces[0].crop.is_empty(), "a crop was invented from nowhere");
assert_eq!(
faces[0].quality, None,
"a quality was invented from nowhere"
);
let _ = std::fs::remove_dir_all(&dir);
}
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -61,7 +61,7 @@ pub use collections::{Collection, CollectionKind, TreeRow};
pub use dedup::{seen_by_content, seen_by_metadata, set_content_hash};
pub use error::CatalogError;
pub use face_shard::{FaceShardStore, SharedFace};
pub use faces::{Calibration, DetectedFace, Face, FaceId, Person, PersonId};
pub use faces::{Calibration, DetectedFace, Face, FaceId, FaceUpdate, Person, PersonId};
pub use jobs::{Job, JobKind, Priority};
pub use keywords::{Coverage, Keyword, KeywordId, SelectionKeyword};
pub use merge::MergeReport;
+59 -5
View File
@@ -991,9 +991,13 @@ fn remote_has_column(tx: &Connection, table: &str, column: &str) -> Result<bool,
/// Remote face row id to local face row id, by photograph and box overlap.
///
/// See [`merge_people_within`] for why a face has no shared identity and this
/// has to be derived. Only faces from the same model are compared: boxes from
/// two different detectors are not the same measurement, and matching across
/// them would attach a judgement to a face nobody looked at.
/// has to be derived. Faces are compared within an *embedder*
/// (`faces::embedder_of`), not within an exact pipeline id: two detectors in
/// front of the same embedder draw boxes around the same faces, and a
/// confirmation made on one device's box is about the face, not the
/// rectangle — the same judgement `faces::record_detections` makes when it
/// carries a confirmation across a re-detection. Keying on the exact id was
/// what let a detector change strand every name on the device that made it.
fn match_faces(tx: &Connection) -> Result<std::collections::HashMap<i64, i64>, CatalogError> {
/// Loose on purpose — "the same face in the frame", not "the same
/// rectangle". The figure `record_detections` uses for the same job.
@@ -1026,7 +1030,8 @@ fn match_faces(tx: &Connection) -> Result<std::collections::HashMap<i64, i64>, C
})?;
for row in rows {
let (file_id, model, boxed) = row?;
local.entry((file_id, model)).or_default().push(boxed);
let embedder = crate::faces::embedder_of(&model).to_string();
local.entry((file_id, embedder)).or_default().push(boxed);
}
}
if local.is_empty() {
@@ -1056,7 +1061,8 @@ fn match_faces(tx: &Connection) -> Result<std::collections::HashMap<i64, i64>, C
for row in rows {
let (remote_id, file_id, model, rbox) = row?;
let Some(candidates) = local.get(&(file_id, model)) else {
let embedder = crate::faces::embedder_of(&model).to_string();
let Some(candidates) = local.get(&(file_id, embedder)) else {
continue;
};
let best = candidates
@@ -1977,6 +1983,54 @@ mod tests {
assert_eq!(person_of(&c, local), Some(("Anna".to_string(), true)));
}
/// The bug this rule exists for: the desktop switched to a stronger
/// detector and confirmed 3,500 faces under the old pipeline id; the
/// tablet held the same faces under the new one, and not one name
/// crossed, because the match demanded the exact id. Same photograph,
/// same box, same embedder — that is the same face.
#[test]
fn a_confirmation_crosses_a_detector_change() {
let c = two_catalogs();
for db in ["main", "remote_cat"] {
add_synced_image(&c, db, 1, 5000);
}
let local = add_face(&c, "main", 7, 1, 0.30);
c.execute(
"UPDATE main.faces SET model_id = 'scrfd_10g+w600k_mbf' WHERE id = ?1",
[local],
)
.unwrap();
let remote = add_face(&c, "remote_cat", 42, 1, 0.31);
add_person(&c, "remote_cat", 3, "u-anna", "Anna", false);
assign(&c, "remote_cat", remote, 3, true);
let report = merge_all(&c).unwrap();
assert_eq!(report.faces_assigned, 1);
assert_eq!(person_of(&c, local), Some(("Anna".to_string(), true)));
}
/// A different embedder is a different space, and a box there is a face
/// nobody here has a vector for.
#[test]
fn a_confirmation_does_not_cross_an_embedder_change() {
let c = two_catalogs();
for db in ["main", "remote_cat"] {
add_synced_image(&c, db, 1, 5000);
}
let local = add_face(&c, "main", 7, 1, 0.30);
c.execute(
"UPDATE main.faces SET model_id = 'scrfd_10g+other_embedder' WHERE id = ?1",
[local],
)
.unwrap();
let remote = add_face(&c, "remote_cat", 42, 1, 0.31);
add_person(&c, "remote_cat", 3, "u-anna", "Anna", false);
assign(&c, "remote_cat", remote, 3, true);
merge_all(&c).unwrap();
assert_eq!(person_of(&c, local), None, "matched across embedders");
}
/// Boxes from two devices are close but not identical. Matching has to be
/// by overlap, not equality, or nothing ever lines up.
#[test]
+1 -7
View File
@@ -301,13 +301,7 @@ fn like_prefix(path: &str) -> String {
}
fn label_code(l: ColourLabel) -> i64 {
match l {
ColourLabel::Red => 1,
ColourLabel::Yellow => 2,
ColourLabel::Green => 3,
ColourLabel::Blue => 4,
ColourLabel::Purple => 5,
}
crate::rating::label_code(l)
}
fn flag_code(f: FlagState) -> i64 {
+428 -13
View File
@@ -30,7 +30,7 @@
use rusqlite::{Connection, OptionalExtension};
use dr_types::{FlagState, ImageId};
use dr_types::{ColourLabel, FlagState, ImageId};
use crate::error::CatalogError;
@@ -63,6 +63,59 @@ impl Judgement {
}
}
/// TRACES: FR-NC-8 | FR-NC-9
/// The default version's uuid for a photograph the server knows by `file_id`.
///
/// # Why this is derived and not generated
///
/// A version's uuid is the identity a cross-device merge keys on. It used to
/// be minted at random per row, and the comment above this function used to
/// say that made it unique — which it did, and that was precisely the bug.
/// Two devices indexing the same library minted *different* uuids for the same
/// photograph, so the sidecar they shared ended up with two `default = 1`
/// blocks, `Version::merge` never saw a matching pair to reconcile, and a
/// rating made on one device was invisible on the other. `crate::merge` has
/// documented the consequence for keywords for as long as it has existed: a
/// uuid-keyed join across two catalogs unions nothing at all.
///
/// `oc:fileid` is the identity that *is* shared. The server assigns it, every
/// client pointed at that library sees the same integer, and it survives a
/// server-side rename and move — the same three properties that made
/// `crate::merge::ASSIGN_BY_FILE_ID` prefer it to a content hash.
///
/// # The layout
///
/// A UUIDv8 (RFC 9562: an application-defined layout) carrying the file id
/// verbatim across the four variable fields, with a fixed tag in the node
/// field saying what minted it. Verbatim rather than hashed so the mapping is
/// injective by construction: two file ids cannot collide, which a truncated
/// hash could, and a uuid read out of a sidecar can be traced back to the file
/// it belongs to by eye.
///
/// Every device computes this identically from the same integer, which is the
/// whole point — there is no negotiation and no first-writer-wins.
pub fn derived_version_uuid(file_id: i64) -> String {
let id = file_id as u64;
format!(
// 32 + 16 + 12 + 4 = 64 bits of file id, then the tag.
"{:08x}-{:04x}-8{:03x}-{:04x}-{:012x}",
(id >> 32) as u32,
(id >> 16) as u16,
(id >> 4) as u16 & 0x0FFF,
// The two high bits are the RFC's variant field and must be `0b10`;
// the remaining fourteen carry the file id's last four bits.
0x8000u16 | ((id as u16 & 0x000F) << 10),
DERIVED_VERSION_TAG,
)
}
/// The node field of a [`derived_version_uuid`], identifying what minted it.
///
/// Fixed and arbitrary. Its only job is to keep a derived uuid from colliding
/// with a randomly minted one and to make it recognisable in a sidecar read by
/// eye — `…-d0c5ec0de001` is visibly not a v4.
const DERIVED_VERSION_TAG: u64 = 0xd0c5_ec0d_e001;
/// Give every image without one a default version.
///
/// Idempotent, and cheap on the common path: the `NOT EXISTS` sub-select is
@@ -72,8 +125,9 @@ impl Judgement {
/// Returns how many were created, so a scan can log the backfill rather than
/// silently doing thousands of inserts.
///
/// The UUID is per row and generated here — it is the merge identity across
/// devices (FR-NC-8), so two images must never share one.
/// The uuid comes from [`derived_version_uuid`] where the server has named the
/// file, so every device computes the same one; only a library with no server
/// behind it falls back to a generated id.
pub fn ensure_default_versions(conn: &Connection) -> Result<usize, CatalogError> {
// One transaction for the batch. A backfill over a 24k-image library is
// 24k inserts, and per-statement commits would make it minutes rather
@@ -93,13 +147,17 @@ pub fn ensure_default_versions(conn: &Connection) -> Result<usize, CatalogError>
/// transaction within a transaction". The same split, for the same reason, as
/// `collections::add_within`.
pub fn ensure_default_versions_within(conn: &Connection) -> Result<usize, CatalogError> {
let ids: Vec<i64> = {
// The remote id travels with the image so the uuid can be derived from it.
// A `LEFT JOIN`, because a library on a folder or a card has no `remote`
// row at all and still needs its versions.
let ids: Vec<(i64, Option<i64>)> = {
let mut stmt = conn.prepare(
"SELECT i.id FROM images i
"SELECT i.id, r.file_id FROM images i
LEFT JOIN remote r ON r.image_id = i.id
WHERE NOT EXISTS (SELECT 1 FROM versions v WHERE v.image_id = i.id)",
)?;
let found = stmt
.query_map([], |r| r.get(0))?
.query_map([], |r| Ok((r.get(0)?, r.get(1)?)))?
.collect::<Result<Vec<_>, _>>()?;
found
};
@@ -112,14 +170,103 @@ pub fn ensure_default_versions_within(conn: &Connection) -> Result<usize, Catalo
"INSERT INTO versions(image_id, uuid, name, is_default, rating, flag)
VALUES (?1, ?2, ?3, 1, 0, 0)",
)?;
for id in &ids {
insert.execute(rusqlite::params![id, new_uuid(), DEFAULT_VERSION_NAME])?;
for (id, file_id) in &ids {
insert.execute(rusqlite::params![
id,
version_uuid(*file_id),
DEFAULT_VERSION_NAME
])?;
}
}
Ok(ids.len())
}
/// The uuid to mint for a new default version.
///
/// Derived from the server's file id where there is one, so two devices agree
/// (FR-NC-8); generated where there is not.
///
/// # What the fallback costs
///
/// A library with no server behind it — a folder, a card — has no identity two
/// devices could both compute, so the split this derivation prevents is still
/// reachable there if that folder is synced by something else. That case is
/// repaired rather than prevented: `Sidecar::fuse_default_versions` folds the
/// rival defaults together the next time either device reads the file.
fn version_uuid(file_id: Option<i64>) -> String {
match file_id {
Some(id) => derived_version_uuid(id),
None => new_uuid(),
}
}
/// TRACES: FR-NC-8 | FR-NC-9
/// Move default versions minted before [`derived_version_uuid`] onto it.
///
/// Every catalog written by an earlier build holds a randomly minted uuid per
/// image, and its peers hold different ones for the same photographs. Deriving
/// the uuid only for *new* rows would leave every image already indexed —
/// which is all of them, on a library anybody has used — writing to the same
/// rival identity it always did.
///
/// Safe to run repeatedly: it selects only rows whose uuid is not already the
/// derived one, so a realigned catalog matches nothing and writes nothing.
///
/// # Why this cannot collide
///
/// `versions.uuid` is `UNIQUE`, and `remote.file_id` has a unique index of its
/// own, so two images cannot derive the same uuid. The one row that could
/// stand in the way is a *virtual copy* (FR-CAT-12) that already holds the
/// target — impossible to mint but not impossible to receive from a merge — so
/// the update is skipped where the target is taken rather than failing the
/// backfill and, with it, the catalog open.
///
/// # The sidecar side is not this function's business
///
/// Moving the catalog's uuid alone would leave the file's default under the
/// old one and the next write would add a rival rather than amend it. What
/// stops that is `library::amend` fusing onto the write's uuid before it looks
/// anything up, which renames the file's default to match. This end and that
/// one have to land together, and they do.
///
/// Returns how many rows moved.
pub fn align_default_version_uuids(conn: &Connection) -> Result<usize, CatalogError> {
// Filtered in SQL rather than in the loop: every derived uuid ends in the
// tag, so a catalog that has already been realigned selects no rows at all
// and this costs one indexed pass instead of twenty-four thousand reads.
let already = format!("%-{DERIVED_VERSION_TAG:012x}");
let stale: Vec<(i64, i64)> = {
let mut stmt = conn.prepare(
"SELECT v.id, r.file_id
FROM versions v
JOIN remote r ON r.image_id = v.image_id
WHERE v.is_default = 1 AND v.uuid NOT LIKE ?1",
)?;
let found = stmt
.query_map([&already], |r| Ok((r.get(0)?, r.get(1)?)))?
.collect::<Result<Vec<_>, _>>()?;
found
};
if stale.is_empty() {
return Ok(0);
}
let tx = conn.unchecked_transaction()?;
let mut moved = 0usize;
{
// `OR IGNORE` covers the taken-target case described above: the row
// keeps the uuid it has, which is the state this build has always
// coped with, rather than aborting the transaction.
let mut update = tx.prepare("UPDATE OR IGNORE versions SET uuid = ?2 WHERE id = ?1")?;
for (row, file_id) in &stale {
moved += update.execute(rusqlite::params![row, derived_version_uuid(*file_id)])?;
}
}
tx.commit()?;
Ok(moved)
}
/// The default version's row id for an image, creating one if it has none.
///
/// Every write path goes through this rather than assuming a version exists.
@@ -128,6 +275,35 @@ pub fn ensure_default_versions_within(conn: &Connection) -> Result<usize, Catalo
/// version pass was interrupted between the image insert and the commit.
/// Failing a rating because of either would be the wrong answer — the user
/// pressed a key and expects a star.
/// TRACES: FR-CAT-13
/// How `versions.label` encodes a colour label, and back.
///
/// One place for both directions, so a label written by the XMP pull and a
/// label queried by the selector cannot drift apart: the query used to hold
/// its own copy of the forward mapping and nothing held the reverse.
pub fn label_code(l: ColourLabel) -> i64 {
match l {
ColourLabel::Red => 1,
ColourLabel::Yellow => 2,
ColourLabel::Green => 3,
ColourLabel::Blue => 4,
ColourLabel::Purple => 5,
}
}
/// The colour a `versions.label` value names, or `None` for NULL and for a
/// code this build does not know.
pub fn label_from_code(code: Option<i64>) -> Option<ColourLabel> {
Some(match code? {
1 => ColourLabel::Red,
2 => ColourLabel::Yellow,
3 => ColourLabel::Green,
4 => ColourLabel::Blue,
5 => ColourLabel::Purple,
_ => return None,
})
}
pub fn default_version_id(conn: &Connection, image: ImageId) -> Result<i64, CatalogError> {
let existing: Option<i64> = conn
.query_row(
@@ -144,10 +320,21 @@ pub fn default_version_id(conn: &Connection, image: ImageId) -> Result<i64, Cata
return Ok(id);
}
// Derived from the server's file id where there is one, so the version
// this mints is the same one the photographer's other device will mint
// (FR-NC-8). A miss here is a library with no server behind it.
let file_id: Option<i64> = conn
.query_row(
"SELECT file_id FROM remote WHERE image_id = ?1",
[image.0 as i64],
|r| r.get(0),
)
.optional()?;
conn.execute(
"INSERT INTO versions(image_id, uuid, name, is_default, rating, flag)
VALUES (?1, ?2, ?3, 1, 0, 0)",
rusqlite::params![image.0 as i64, new_uuid(), DEFAULT_VERSION_NAME],
rusqlite::params![image.0 as i64, version_uuid(file_id), DEFAULT_VERSION_NAME],
)?;
Ok(conn.last_insert_rowid())
}
@@ -356,15 +543,22 @@ fn flag_from_code(v: i64) -> FlagState {
}
}
/// A version UUID.
/// A generated version UUID, for a photograph no server has named.
///
/// Hand-rolled rather than pulling in the `uuid` crate for one function — the
/// same reasoning as the date maths in `library_ui`. This needs to be unique
/// across devices, not cryptographically unguessable: it keys a merge, and an
/// attacker who can write to the sidecar has already won.
/// same reasoning as the date maths in `library_ui`. It needs to be unique,
/// not cryptographically unguessable: it keys a merge, and an attacker who can
/// write to the sidecar has already won.
///
/// Seeded from the system clock and a per-process counter, so two versions
/// created inside the same nanosecond tick still differ.
///
/// **Unique is not the same as agreed**, which is the distinction that cost a
/// photographer a day of culling. Two devices calling this for the same
/// photograph get two different answers, and a merge keyed on the result then
/// has no pair to reconcile. Anything with a `file_id` behind it must use
/// [`derived_version_uuid`]; this is the fallback for libraries that have no
/// server to supply one.
fn new_uuid() -> String {
use std::sync::atomic::{AtomicU64, Ordering};
static COUNTER: AtomicU64 = AtomicU64::new(0);
@@ -731,3 +925,224 @@ mod tests {
assert_eq!(cat.count(&unrated, 0).unwrap(), 4);
}
}
/// TRACES: FR-NC-8 | FR-NC-9
/// The identity two devices have to agree on without talking to each other.
#[cfg(test)]
mod derived_identity {
use super::*;
use crate::Catalog;
/// A catalog whose images the server has named, as a remote scan leaves it.
fn with_remote_images(file_ids: &[i64]) -> Catalog {
let cat = Catalog::in_memory().unwrap();
let c = cat.connection();
c.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'remote', 'lib')",
[],
)
.unwrap();
for (i, file_id) in file_ids.iter().enumerate() {
c.execute(
"INSERT INTO images(root_id, source_ref, added_at) VALUES (1, ?1, 0)",
[format!("img{i:03}.CR3")],
)
.unwrap();
let image = c.last_insert_rowid();
c.execute(
"INSERT INTO remote(image_id, file_id) VALUES (?1, ?2)",
rusqlite::params![image, file_id],
)
.unwrap();
}
cat
}
fn default_uuids(cat: &Catalog) -> Vec<String> {
let mut stmt = cat
.connection()
.prepare("SELECT uuid FROM versions WHERE is_default = 1 ORDER BY image_id")
.unwrap();
stmt.query_map([], |r| r.get(0))
.unwrap()
.map(Result::unwrap)
.collect()
}
/// The whole point: the same photograph, indexed independently on two
/// devices, gets one identity. This used to be two.
#[test]
fn two_devices_derive_the_same_uuid_for_one_photograph() {
let laptop = with_remote_images(&[4_812]);
let tablet = with_remote_images(&[4_812]);
ensure_default_versions(laptop.connection()).unwrap();
ensure_default_versions(tablet.connection()).unwrap();
assert_eq!(default_uuids(&laptop), default_uuids(&tablet));
}
/// And different photographs must still be told apart — the property the
/// random uuid did have, which this must not give up to gain agreement.
#[test]
fn different_photographs_keep_different_uuids() {
let cat = with_remote_images(&[1, 2, 3, 0x7FFF_FFFF_FFFF_FFFF]);
ensure_default_versions(cat.connection()).unwrap();
let mut uuids = default_uuids(&cat);
let before = uuids.len();
uuids.sort();
uuids.dedup();
assert_eq!(uuids.len(), before, "two photographs share an identity");
}
/// The file id has to survive the layout intact, or two ids that differ
/// only in the bits it drops would collide.
#[test]
fn the_whole_file_id_is_carried() {
// A pair differing only in the low four bits, and a pair differing
// only in the high thirty-two — the two places a sloppy layout loses
// information.
assert_ne!(derived_version_uuid(0x10), derived_version_uuid(0x1F));
assert_ne!(
derived_version_uuid(0x0000_0001_0000_0000),
derived_version_uuid(0x0000_0002_0000_0000)
);
assert_ne!(derived_version_uuid(0), derived_version_uuid(-1));
}
/// Well-formed, and recognisably not a generated one.
#[test]
fn a_derived_uuid_is_a_well_formed_v8() {
let uuid = derived_version_uuid(4_812);
let fields: Vec<&str> = uuid.split('-').collect();
assert_eq!(fields.len(), 5);
assert_eq!(
fields.iter().map(|f| f.len()).collect::<Vec<_>>(),
vec![8, 4, 4, 4, 12]
);
assert!(fields[2].starts_with('8'), "version nibble: {uuid}");
// The RFC's variant field is the two high bits of the fourth group,
// and must read `0b10` — so the first hex digit is 8, 9, a or b.
assert!(
matches!(fields[3].as_bytes()[0], b'8' | b'9' | b'a' | b'b'),
"variant: {uuid}"
);
assert!(uuid.ends_with("d0c5ec0de001"), "tag: {uuid}");
}
/// A library with no server behind it has no shared identity to derive,
/// and must still get a version rather than failing.
#[test]
fn a_library_with_no_server_still_gets_its_versions() {
let cat = Catalog::in_memory().unwrap();
let c = cat.connection();
c.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'local', 'lib')",
[],
)
.unwrap();
c.execute(
"INSERT INTO images(root_id, source_ref, added_at) VALUES (1, 'a.CR3', 0)",
[],
)
.unwrap();
assert_eq!(ensure_default_versions(c).unwrap(), 1);
assert_eq!(default_uuids(&cat).len(), 1);
}
/// The repair. A catalog written by an earlier build holds randomly minted
/// uuids, and leaving them there would mean every image already indexed —
/// which is all of them — kept writing to its own rival identity.
#[test]
fn a_catalog_from_an_earlier_build_is_realigned() {
let cat = with_remote_images(&[4_812, 4_813]);
let c = cat.connection();
// As the old code left it.
for (i, image) in [1i64, 2].iter().enumerate() {
c.execute(
"INSERT INTO versions(image_id, uuid, name, is_default, rating, flag)
VALUES (?1, ?2, 'Default', 1, ?3, 0)",
rusqlite::params![image, format!("random-{i}"), (i + 1) as i64],
)
.unwrap();
}
assert_eq!(align_default_version_uuids(c).unwrap(), 2);
assert_eq!(
default_uuids(&cat),
vec![derived_version_uuid(4_812), derived_version_uuid(4_813)]
);
// The judgement travels with the row — a realignment that dropped the
// ratings would be a worse bug than the one it fixes.
let ratings: Vec<i64> = {
let mut stmt = c
.prepare("SELECT rating FROM versions ORDER BY image_id")
.unwrap();
let v = stmt
.query_map([], |r| r.get(0))
.unwrap()
.map(Result::unwrap)
.collect();
v
};
assert_eq!(ratings, vec![1, 2]);
}
/// Runs on every catalog open, so a second pass must select nothing and
/// write nothing.
#[test]
fn realigning_twice_changes_nothing_the_second_time() {
let cat = with_remote_images(&[4_812]);
ensure_default_versions(cat.connection()).unwrap();
assert_eq!(
align_default_version_uuids(cat.connection()).unwrap(),
0,
"a freshly derived catalog must match nothing"
);
let before = default_uuids(&cat);
align_default_version_uuids(cat.connection()).unwrap();
assert_eq!(default_uuids(&cat), before);
}
/// A virtual copy (FR-CAT-12) that already holds the target uuid must not
/// take the backfill — and with it the catalog open — down with it.
#[test]
fn a_taken_target_leaves_the_row_where_it_is() {
let cat = with_remote_images(&[4_812]);
let c = cat.connection();
c.execute(
"INSERT INTO versions(image_id, uuid, name, is_default, rating, flag)
VALUES (1, 'random', 'Default', 1, 0, 0)",
[],
)
.unwrap();
c.execute(
"INSERT INTO versions(image_id, uuid, name, is_default, rating, flag)
VALUES (1, ?1, 'For print', 0, 0, 0)",
[derived_version_uuid(4_812)],
)
.unwrap();
assert_eq!(align_default_version_uuids(c).unwrap(), 0);
assert_eq!(default_uuids(&cat), vec!["random".to_string()]);
}
/// The backfill is what actually runs this, so it has to be wired in.
#[test]
fn opening_a_catalog_realigns_it() {
let cat = with_remote_images(&[4_812]);
cat.connection()
.execute(
"INSERT INTO versions(image_id, uuid, name, is_default, rating, flag)
VALUES (1, 'random', 'Default', 1, 3, 0)",
[],
)
.unwrap();
crate::schema::backfill(cat.connection()).unwrap();
assert_eq!(default_uuids(&cat), vec![derived_version_uuid(4_812)]);
}
}
+91
View File
@@ -198,6 +198,54 @@ pub fn backup_before_migration(conn: &Connection, catalog: &Path) -> Result<(),
Ok(())
}
/// How long a catalog may go without a backup before the next opportunity
/// takes one.
///
/// A day. The catalog is an index, so what a backup protects is the day's
/// worth of collection and people edits the sidecars do not hold — and a
/// second copy of a 130 MB file per launch would be a cost with nothing to
/// show for it when the user launches four times in an afternoon.
pub const BACKUP_EVERY: i64 = 24 * 60 * 60;
/// Whether [`BACKUP_EVERY`] has passed since the newest backup, or there is
/// none.
///
/// Read from the filenames, like [`backups`], so a restored or copied backup
/// directory answers the same way it did on the machine it came from.
pub fn backup_due(catalog: &Path) -> bool {
match backups(catalog).first() {
Some(newest) => now() - newest.taken_at >= BACKUP_EVERY,
None => true,
}
}
/// TRACES: NFR-R2
/// Take the scheduled backup, if one is due. Returns the file written, or
/// `None` when the newest is recent enough.
///
/// The scheduled half of NFR-R2 — the migration half is
/// [`backup_before_migration`]. "On a schedule" for an application that runs
/// when the user opens it means "at the next chance after a day has passed",
/// and the chance the caller picks is the end of a library sweep: the
/// catalog is quiet, the work is already off the UI thread, and it is the
/// moment a day's edits have just been consolidated.
///
/// A brand-new catalog with no images is not backed up: there is nothing in
/// it yet that a rescan would not rebuild, and the first backup would only be
/// a copy of an empty schema.
pub fn backup_if_due(conn: &Connection, catalog: &Path) -> Result<Option<PathBuf>, CatalogError> {
if !backup_due(catalog) {
return Ok(None);
}
let images: i64 = conn.query_row("SELECT count(*) FROM images", [], |r| r.get(0))?;
if images == 0 {
return Ok(None);
}
let path = backup(conn, catalog)?;
log::info!("scheduled backup of the catalog to {}", path.display());
Ok(Some(path))
}
/// The backups available for `catalog`, newest first.
///
/// Never fails: an unreadable or absent backup directory means there are no
@@ -386,6 +434,49 @@ mod tests {
base
}
#[test]
fn a_scheduled_backup_is_taken_once_a_day_and_not_more() {
let dir = tempdir("scheduled");
let path = dir.join("catalog.sqlite");
fixture(&path, 3);
let cat = Catalog::open(&path).unwrap();
// Nothing yet: due.
assert!(backup_due(&path));
let first = backup_if_due(cat.connection(), &path).unwrap();
assert!(first.is_some(), "the first opportunity takes one");
// Taken just now: not due, and a second call does nothing.
assert!(!backup_due(&path));
assert_eq!(backup_if_due(cat.connection(), &path).unwrap(), None);
assert_eq!(backups(&path).len(), 1);
// Age the one backup past the interval by renaming it, since the
// timestamp is read from the name. Now it is due again.
let old = first.unwrap();
let aged = old
.parent()
.unwrap()
.join(format!("catalog-{}.sqlite", now() - BACKUP_EVERY - 1));
std::fs::rename(&old, &aged).unwrap();
assert!(backup_due(&path));
assert!(backup_if_due(cat.connection(), &path).unwrap().is_some());
assert_eq!(backups(&path).len(), 2);
let _ = std::fs::remove_dir_all(&dir);
}
#[test]
fn an_empty_catalog_is_not_worth_backing_up() {
let dir = tempdir("empty");
let path = dir.join("catalog.sqlite");
let cat = Catalog::open(&path).unwrap();
assert!(backup_due(&path), "due in principle");
assert_eq!(backup_if_due(cat.connection(), &path).unwrap(), None);
assert!(backups(&path).is_empty());
let _ = std::fs::remove_dir_all(&dir);
}
/// A catalog on disk with enough rows to span several pages, closed.
///
/// Closed matters: WAL means the rows are in `catalog.sqlite-wal` until
+548 -7
View File
@@ -15,7 +15,7 @@ use rusqlite::Connection;
use crate::error::CatalogError;
/// Schema version this build writes and understands.
pub const SCHEMA_VERSION: i64 = 11;
pub const SCHEMA_VERSION: i64 = 18;
/// Apply migrations up to [`SCHEMA_VERSION`].
///
@@ -105,9 +105,99 @@ pub fn migrate(conn: &Connection) -> Result<i64, CatalogError> {
tx.commit()?;
}
if from < 12 {
let tx = conn.unchecked_transaction()?;
tx.execute_batch(V12)?;
tx.pragma_update(None, "user_version", 12)?;
tx.commit()?;
}
if from < 13 {
let tx = conn.unchecked_transaction()?;
tx.execute_batch(V13)?;
tx.pragma_update(None, "user_version", 13)?;
tx.commit()?;
}
if from < 14 {
let tx = conn.unchecked_transaction()?;
// `ALTER TABLE ... ADD COLUMN` has no `IF NOT EXISTS`, and NFR-R5
// wants this re-enterable: a catalog whose `user_version` was rewound
// by a rollback already has the column, and would otherwise fail its
// next open on it.
let has_quality: bool = tx
.prepare("SELECT 1 FROM pragma_table_info('faces') WHERE name = 'quality'")?
.exists([])?;
if !has_quality {
tx.execute_batch("ALTER TABLE faces ADD COLUMN quality REAL;")?;
}
tx.execute_batch(V14)?;
tx.pragma_update(None, "user_version", 14)?;
tx.commit()?;
}
if from < 15 {
let tx = conn.unchecked_transaction()?;
tx.execute_batch(V15)?;
tx.pragma_update(None, "user_version", 15)?;
tx.commit()?;
}
if from < 16 {
let tx = conn.unchecked_transaction()?;
// Guarded like V14's column, and for the same reason: `ALTER TABLE
// ... ADD COLUMN` has no `IF NOT EXISTS`, and this step must be
// re-enterable (NFR-R5).
for column in EYE_COLUMNS {
let present: bool = tx
.prepare("SELECT 1 FROM pragma_table_info('faces') WHERE name = ?1")?
.exists([column])?;
if !present {
tx.execute_batch(&format!("ALTER TABLE faces ADD COLUMN {column} REAL;"))?;
}
}
tx.pragma_update(None, "user_version", 16)?;
tx.commit()?;
}
if from < 17 {
let tx = conn.unchecked_transaction()?;
tx.execute_batch(V17)?;
tx.pragma_update(None, "user_version", 17)?;
tx.commit()?;
}
if from < 18 {
let tx = conn.unchecked_transaction()?;
// Guarded like V14's and V16's columns: ALTER has no IF NOT EXISTS
// and the step must be re-enterable (NFR-R5).
let present: bool = tx
.prepare("SELECT 1 FROM pragma_table_info('faces') WHERE name = 'landmarks_dense'")?
.exists([])?;
if !present {
tx.execute_batch("ALTER TABLE faces ADD COLUMN landmarks_dense BLOB;")?;
}
tx.pragma_update(None, "user_version", 18)?;
tx.commit()?;
}
Ok(from)
}
/// The seven columns V16 adds to `faces`, in the order the readers name them.
///
/// Named once because three places have to agree on them: this migration,
/// [`for_attached`], and the face shard's own catch-up (`face_shard`).
pub const EYE_COLUMNS: [&str; 7] = [
"eye_right",
"eye_right_px",
"eye_right_sharp",
"eye_left",
"eye_left_px",
"eye_left_sharp",
"sunglasses",
];
/// Recompute columns a migration added, for rows that predate it.
///
/// A migration adds a column with a default; it cannot know what the value
@@ -132,15 +222,27 @@ pub fn backfill(conn: &Connection) -> Result<Vec<(&'static str, usize)>, Catalog
// v3: every image needs a default version to carry its rating and flag.
// Libraries scanned before ratings existed have images and no versions at
// all, so there was nowhere for a judgement to go — see
// [`crate::rating`]. Backfilled rather than migrated in SQL because the
// UUID per row is the cross-device merge identity and must be generated,
// not derived.
// all, so there was nowhere for a judgement to go — see [`crate::rating`].
let n = crate::rating::ensure_default_versions(conn)?;
if n > 0 {
out.push(("default_versions", n));
}
// TRACES: FR-NC-8 | FR-NC-9
// The uuid on those rows is the cross-device merge identity, and it used
// to be generated rather than derived. This comment said so, and said it
// as though generating it were the point — it was the bug. Two devices
// minted different uuids for one photograph, so the sidecar they shared
// grew a `default = 1` block each and neither ever saw the other's work.
//
// Runs after the pass above so a row created a moment ago is already
// derived and matches nothing here. Ordering the other way would be
// correct too, just wasteful.
let n = crate::rating::align_default_version_uuids(conn)?;
if n > 0 {
out.push(("derived_version_uuids", n));
}
// v6: a vocabulary row for every word some image already carries.
//
// Three ways a catalog arrives holding assignments with no term behind
@@ -158,11 +260,38 @@ pub fn backfill(conn: &Connection) -> Result<Vec<(&'static str, usize)>, Catalog
Ok(out)
}
/// How long a connection waits for a writer to finish before giving up.
///
/// TRACES: NFR-R1
/// SQLite's default is **zero** — the loser of a race gets `SQLITE_BUSY` at
/// once rather than a turn — and WAL does not change that for two writers. One
/// writer and many readers is the case WAL makes free; this is the other one,
/// and this application has it constantly: the face sweep commits a batch while
/// reclustering reads, the derived sync imports shards while the sweep writes.
///
/// Without a timeout that contention was *lost work*, not a retry. A face
/// sweep that had already paid for the detection and the embedding — the
/// expensive part, seconds per image — threw the result away on
/// `storing faces for 214: database is locked` and moved on, and both the
/// desktop and the tablet logged runs of those on consecutive images.
///
/// Ten seconds, matching the figure the job runner's tests already use for the
/// same reason. It is far longer than any transaction here (a sweep batch is
/// sub-second; the slowest is a WAL checkpoint of a 130 MB catalog), so in
/// practice it is a bound on pathology rather than a wait anyone sits through.
/// The tension with NFR-P9 is real but one-sided: a query on the UI thread
/// would rather wait for its turn than fail, because the failure is what the
/// user sees as "cannot open catalog".
const BUSY_TIMEOUT: std::time::Duration = std::time::Duration::from_secs(10);
/// Connection setup applied on every open, migration or not.
///
/// WAL is required by NFR-R1: it survives power loss without corruption, and
/// it lets a background job write while the grid reads.
pub fn configure(conn: &Connection) -> Result<(), CatalogError> {
// Before the pragmas, so that a connection racing a migration waits for it
// rather than failing on the first statement it tries.
conn.busy_timeout(BUSY_TIMEOUT)?;
conn.pragma_update(None, "journal_mode", "WAL")?;
// NORMAL rather than FULL: with WAL this is durable across process death
// (which is what FR-PLAT-AND-3 cares about) and only risks the last
@@ -225,7 +354,16 @@ pub fn for_attached(schema_name: &str) -> String {
format!(
"{}\n{}\n{}\n\
ALTER TABLE {schema_name}.people ADD COLUMN ignored INTEGER NOT NULL DEFAULT 0;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN crop BLOB;",
ALTER TABLE {schema_name}.faces ADD COLUMN crop BLOB;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN quality REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN eye_right REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN eye_right_px REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN eye_right_sharp REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN eye_left REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN eye_left_px REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN eye_left_sharp REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN sunglasses REAL;\n\
ALTER TABLE {schema_name}.faces ADD COLUMN landmarks_dense BLOB;",
rewrite_for_attached(V1, schema_name),
rewrite_for_attached(V6, schema_name),
rewrite_for_attached(V8, schema_name),
@@ -476,6 +614,224 @@ CREATE TABLE burst_expanded (
);
"#;
const V12: &str = r#"
-- TRACES: FR-CULL-8
-- Forget the runs that were made against a proxy too small to find a face on.
--
-- Detection used to accept any proxy, and one of the two sweeps detected on
-- the stored 1024px tier. On the reference library that produced 0.078 faces
-- per image against 1.82 for the same photographs at 2048 or better -- and
-- every one of those runs left a `face_index` row behind saying the image had
-- been examined. That row is what makes the damage permanent: the work list is
-- "images with no row", so a photograph examined badly is indistinguishable
-- from one examined well, and is never offered to a later pass.
--
-- Deleting the marker is the whole repair, and it is deliberately not a
-- deletion of anything else. The `faces` rows those runs found stay exactly
-- where they are and keep drawing the People screen until a better pass
-- replaces them, and `record_detections` carries the user's confirmed names
-- across that replacement by box overlap. So this costs a re-fetch of the
-- affected images and loses no work the user has done.
--
-- The threshold is written out rather than taken from `dr_face::MIN_CROP_EDGE`
-- on purpose. A migration has to keep meaning what it meant on the day it ran;
-- binding it to a constant someone may raise later would silently change what
-- an old catalog gets migrated to.
DELETE FROM face_index WHERE source_edge <= 1024;
"#;
const V13: &str = r#"
-- TRACES: FR-CAT-8 | FR-NC-9
-- Which sidecars this device has read, and at what ETag.
--
-- The sidecar is the authoritative store for a rating and an edit, and until
-- this table existed nothing ever read one back into the catalog: judgements
-- travelled outward only. A cull done on a tablet reached the server and
-- stopped there, because the scan indexes photographs, the derived sync moves
-- thumbnails and collections, and the one reader that existed ran when a single
-- photograph was opened in develop and fed only the develop graph. The grid
-- draws `versions.rating`, so another device's afternoon of culling was
-- invisible on this one -- permanently, by every path the app had.
--
-- What this holds is the ETag, not the content. It is the record of what has
-- already been taken in, so a pull fetches only what changed: `dr_sync::scan`
-- reports every sidecar it saw in listings it was making anyway, and this
-- decides which of them are worth a GET.
--
-- Keyed on the sidecar's own remote path rather than on an image id. One
-- sidecar can describe two images -- a RAW and the JPEG beside it are one
-- photograph (FR-CAT-11) and share a document -- and a path is what the scan
-- reports and what a fetch addresses, so keying on anything else would mean
-- deriving one from the other in two places.
--
-- Rebuildable like the rest of the catalog: losing this table costs one pass
-- that re-reads every sidecar and reaches exactly the same state.
--
-- `IF NOT EXISTS` because NFR-R5 asks for migrations that are idempotent on
-- retry, and this one can genuinely be re-entered: a catalog whose
-- `user_version` was rewound -- by a rollback to an older build, or by a
-- recovery -- would otherwise fail its next open on a table it already has.
CREATE TABLE IF NOT EXISTS sidecars (
root_id INTEGER NOT NULL REFERENCES roots(id) ON DELETE CASCADE,
path TEXT NOT NULL,
etag TEXT,
-- Unix seconds, for diagnosing a pull that is not making progress.
read_at INTEGER NOT NULL DEFAULT 0,
PRIMARY KEY(root_id, path)
);
"#;
const V14: &str = r#"
-- TRACES: FR-CULL-9 | FR-CULL-10
-- How recognisable the model found each face, and a second look at the faces
-- it was never asked about.
--
-- The embedder's raw output has a length, and the length is a quality
-- reading: it grows with how much of a face the model could make out, and a
-- blur, an occlusion or a hard profile comes out short (dr_face::embedding,
-- `MIN_GALLERY_QUALITY`). Normalising threw it away. A short vector sits
-- near the middle of the sphere and matches a little of everyone, which is
-- how one bad crop bridges two people in a grouping pass -- so a face below
-- the floor is compared against the others and never compared *against*.
--
-- Nullable, and NULL means "never measured": every face indexed before this
-- version stored the unit vector, whose length is one whatever the crop was.
-- A face with no reading is admitted to the gallery, because a rule that
-- cannot be checked should admit rather than exclude -- but it is also a
-- face this rule is not yet protecting anyone from, and the only way to
-- measure it is to embed it again.
--
-- The `face-quality` repair is what does that (`dr_ui::repairs`, once the
-- sweep's measuring pass): it lists every face with no reading, and each is
-- embedded again from the native render with the landmarks it already has,
-- the raw vector written over the old one (`record_updates`) and nothing
-- else touched -- not the id, not the box, not who the user said it was.
-- The faces keep drawing the People screen throughout.
--
-- The run markers of those images are forgotten too, exactly as V12 forgot
-- the runs made against too small a proxy. The build this shipped in had no
-- measuring pass yet, and a marker is the one thing that stops a face ever
-- being looked at again; with the repair in place, detection leaves an
-- image holding this embedder's faces to it rather than detecting from
-- scratch, so the deletion costs nothing -- and an image that was examined
-- and found empty keeps its marker, since there is nothing on it to measure.
--
-- The cost is a re-fetch of every image with a face on it, on the next pass
-- the user starts. That is a whole-library transfer (FR-NC-6), and it starts
-- when they say so, not here.
--
-- From this version the `embedding` blob is the **raw** model output rather
-- than the unit vector V8 describes -- the length is the quality, and a store
-- that kept only the direction had thrown it away. Readers re-normalise on
-- load, so a unit blob from before and a raw blob from now compare alike;
-- `quality` is that length kept beside the blob for the readers that never
-- load the vector, and NULL rather than 1.0 for the old rows, because a unit
-- vector reads as a length of one and one is not "unmeasured".
--
-- The column itself is added in `migrate`, guarded, because ALTER has no
-- IF NOT EXISTS and this step has to be re-enterable (NFR-R5).
DELETE FROM face_index
WHERE EXISTS (SELECT 1 FROM faces f
WHERE f.image_id = face_index.image_id
AND f.model_id = face_index.model_id);
"#;
const V15: &str = r#"
-- TRACES: FR-CAT-13
-- Where a standard XMP sidecar and the catalog disagree.
--
-- An `.xmp` beside a photograph is read on the same pull as DarkRoom's own
-- sidecar, and reconciled field by field (`dr_xmp::reconcile`): keywords
-- union, and a rating, label or caption is taken only where the catalog holds
-- none. That rule is the safe one and it is not always the right one -- a
-- rating changed in Lightroom after it was changed here is a genuine
-- disagreement, and a standard XMP carries no revision to settle it by. So
-- the disagreement is written here instead of being resolved, and the
-- requirement's "a metadata reload offered" is a row in this table with a
-- button in front of it: the reload re-reads the file with the sidecar
-- winning, and deletes the row.
--
-- Keyed on the sidecar's path like `sidecars` is, and for the same reason: a
-- path is what the scan reports, what a fetch addresses, and what the ETag
-- that noticed the change belongs to. `fields` is the disagreeing fields as
-- `dr_xmp` names them, space-separated, for the line the settings page shows.
--
-- Rebuildable: the next pull that sees a changed ETag writes the row again.
CREATE TABLE IF NOT EXISTS xmp_conflicts (
root_id INTEGER NOT NULL REFERENCES roots(id) ON DELETE CASCADE,
path TEXT NOT NULL,
fields TEXT NOT NULL,
seen_at INTEGER NOT NULL DEFAULT 0,
PRIMARY KEY(root_id, path)
);
"#;
// V18 -- TRACES: FR-CULL-8a | FR-CULL-12
//
// The 106 dense landmarks the eye pass reads its eye boxes from, kept beside
// the reading as `dr_face::Landmarks::to_packed_bytes`: 106 x (x, y) as
// 16-bit fixed point over the frame, 424 bytes a face, a seventh of a
// pixel on a 6000-pixel frame. Derived data under FR-CULL-12 -- rebuilt by
// re-reading, never in a sidecar -- and stored for the same reason the
// embedding is: it cost a fetch of the original and a model run, and the
// next per-face pass (head pose, expression) should not have to pay either
// again. NULL where the face was never read.
//
// Added in `migrate`, guarded, like every ALTER here (NFR-R5).
const V17: &str = r#"
-- TRACES: FR-CULL-8a | FR-CULL-13 | NFR-P9
-- The eyes-open filter's index, and a lesson about where a column lands.
--
-- The people filter is a correlated EXISTS over `faces` per image, and it
-- was fast because `faces_image` *covers* it: the subquery never touched a
-- row. Reading V16's seven eye columns in the same subquery did touch the
-- row -- and `ALTER TABLE ADD COLUMN` puts a column at the end of the
-- record, after the 1 KB embedding and the ~5 KB crop, so every check
-- dragged six kilobytes off disk to reach seven floats. Measured on the
-- reference library: 24 seconds for one count, thirteen of them system
-- time. With this index the same count takes five milliseconds, because
-- the subquery is served from the index again and never reads a row.
--
-- The columns are listed in EYE_COLUMNS' order behind `image_id`, which is
-- the key the subquery searches on. Nothing else changed in V17; a catalog
-- already at V16 needs only this.
CREATE INDEX IF NOT EXISTS faces_eyes ON faces(
image_id, eye_right, eye_right_px, eye_right_sharp,
eye_left, eye_left_px, eye_left_sharp, sunglasses
);
"#;
// V16 -- TRACES: FR-CULL-8a
//
// What each face's eyes are doing: for each eye P(open), the source pixels
// across its box and the sharpness of the patch the classifier saw; and
// P(sunglasses) for the head. Seven numbers rather than a verdict, because
// the verdict is a rule with thresholds in it (dr_face::eyes::EyeReading::
// state) and a rule belongs in code that can be changed, not in rows that
// would have to be re-measured.
//
// The pixels and the sharpness are what stop a smear reading as a blink: an
// eye too small or too soft to read is not asked, and a face with no
// readable eye is "unclear", which no filter drops. Sunglasses are a column
// of their own for the same kind of reason — the eye classifier answers
// confidently over dark glass, and its answer means nothing there. A filter
// for "eyes open" reads all seven.
//
// NULL means "never measured" -- a face indexed before this version, or on a
// device without the eye models -- and a NULL is left alone by every filter
// that reads these, so an old library does not empty its grid the moment the
// chip is pressed. The sweep's measuring pass fills them in, from the native
// render, with the landmarks already stored: the same pass V14 built for the
// embedding's length, extended to ask the eye models too. No run marker is
// forgotten here, for the reason V14's note gives -- the measuring pass
// finds its own work by the NULL, and deleting markers would only put the
// detector back over images it has finished with.
//
// The columns are added in `migrate`, guarded, because ALTER has no IF NOT
// EXISTS and the step has to be re-enterable (NFR-R5). Their names are
// `EYE_COLUMNS`.
const V9: &str = r#"
-- TRACES: FR-CULL-8
-- A record that face detection has *run* on an image, distinct from what it
@@ -555,7 +911,7 @@ CREATE TABLE faces (
x REAL NOT NULL, y REAL NOT NULL, w REAL NOT NULL, h REAL NOT NULL,
landmarks BLOB NOT NULL, -- 5 x (x, y) f32, normalised likewise
detector_confidence REAL NOT NULL,
embedding BLOB NOT NULL, -- 512 x f16, L2-normalised
embedding BLOB NOT NULL, -- 512 x f16; unit length until V14, raw since
-- Source pixels across the aligned 112x112 crop (docs/faces.md §7).
--
-- Not cosmetic: it is the honest quality signal for the UI, a feature in
@@ -945,6 +1301,60 @@ CREATE INDEX jobs_ready ON jobs(state, priority DESC, not_before);
#[cfg(test)]
mod tests {
#[test]
fn a_writer_waits_for_its_turn_rather_than_losing_its_work() {
// The failure this exists for: a face sweep that had already paid for
// the detection and the embedding threw the result away on
// "database is locked" and moved on. WAL does not help here — it makes
// one writer and many readers free, and this is two writers.
let dir = std::env::temp_dir().join(format!(
"dr-busy-{}-{:?}",
std::process::id(),
std::thread::current().id()
));
let _ = std::fs::remove_dir_all(&dir);
std::fs::create_dir_all(&dir).unwrap();
let path = dir.join("catalog.sqlite");
let held = rusqlite::Connection::open(&path).unwrap();
configure(&held).unwrap();
migrate(&held).unwrap();
let other = rusqlite::Connection::open(&path).unwrap();
configure(&other).unwrap();
// Every connection carries the timeout, which is what makes the wait
// below a wait rather than an immediate error.
let timeout: i64 = other
.query_row("PRAGMA busy_timeout", [], |r| r.get(0))
.unwrap();
assert_eq!(timeout, BUSY_TIMEOUT.as_millis() as i64);
// A writer holds the database; the other one must still get its turn
// once the first commits, rather than failing at the moment it asks.
let writing = held.unchecked_transaction().unwrap();
held.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'local', 'lib')",
[],
)
.unwrap();
let handle = std::thread::spawn(move || {
other.execute(
"INSERT INTO roots(id, kind, label) VALUES (2, 'local', 'two')",
[],
)
});
std::thread::sleep(std::time::Duration::from_millis(150));
writing.commit().unwrap();
assert!(
handle.join().unwrap().is_ok(),
"the second writer waited and then wrote, rather than erroring"
);
let _ = std::fs::remove_dir_all(&dir);
}
use super::*;
fn mem() -> Connection {
@@ -1232,6 +1642,35 @@ mod tests {
assert_eq!(migrate(&c).unwrap(), SCHEMA_VERSION);
}
/// V16 adds its columns guarded, so a catalog whose version was rewound
/// after the columns landed — the rollback NFR-R5 contemplates — migrates
/// again rather than failing on "duplicate column".
#[test]
fn the_eye_columns_survive_a_rewound_version() {
let c = mem();
migrate(&c).unwrap();
for column in EYE_COLUMNS {
let present: bool = c
.prepare("SELECT 1 FROM pragma_table_info('faces') WHERE name = ?1")
.unwrap()
.exists([column])
.unwrap();
assert!(present, "{column} missing after migration");
}
c.pragma_update(None, "user_version", 15).unwrap();
assert_eq!(migrate(&c).unwrap(), 15);
let indexed: bool = c
.prepare("SELECT 1 FROM sqlite_master WHERE type = 'index' AND name = 'faces_eyes'")
.unwrap()
.exists([])
.unwrap();
assert!(indexed, "V17's covering index is there");
let v: i64 = c
.query_row("PRAGMA user_version", [], |r| r.get(0))
.unwrap();
assert_eq!(v, SCHEMA_VERSION);
}
#[test]
fn refuses_a_catalog_from_a_newer_build() {
let c = mem();
@@ -1268,6 +1707,108 @@ mod tests {
assert_eq!(n, 0, "images must not outlive their root");
}
#[test]
fn v12_forgets_runs_made_on_a_proxy_too_small_to_see_a_face() {
let c = mem();
// Migrate to 11, then seed the state V12 exists to repair: markers
// written at the 1024 store tier beside ones written on a real
// preview.
c.pragma_update(None, "user_version", 0).unwrap();
migrate(&c).unwrap();
c.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'local', 'test')",
[],
)
.unwrap();
c.execute(
"INSERT INTO images(id, root_id, source_ref, added_at)
VALUES (1,1,'a',0),(2,1,'b',0),(3,1,'c',0),(4,1,'d',0)",
[],
)
.unwrap();
for (image, edge) in [(1, 896), (2, 1024), (3, 1025), (4, 2560)] {
c.execute(
"INSERT INTO face_index(image_id, model_id, indexed_at, faces_found, source_edge)
VALUES (?1, 'm', 0, 0, ?2)",
rusqlite::params![image, edge],
)
.unwrap();
}
c.pragma_update(None, "user_version", 11).unwrap();
migrate(&c).unwrap();
let kept: Vec<i64> = c
.prepare("SELECT image_id FROM face_index ORDER BY image_id")
.unwrap()
.query_map([], |r| r.get(0))
.unwrap()
.map(Result::unwrap)
.collect();
// 1024 goes: it is exactly ThumbSize::Large, the tier that produced
// the bad runs. 1025 stays, or the floor and the repair disagree
// about the same boundary.
assert_eq!(kept, vec![3, 4]);
}
#[test]
fn v14_forgets_runs_that_found_faces_but_never_measured_them() {
let c = mem();
c.pragma_update(None, "user_version", 0).unwrap();
migrate(&c).unwrap();
c.execute(
"INSERT INTO roots(id, kind, label) VALUES (1, 'local', 'test')",
[],
)
.unwrap();
c.execute(
"INSERT INTO images(id, root_id, source_ref, added_at)
VALUES (1,1,'a',0),(2,1,'b',0),(3,1,'c',0)",
[],
)
.unwrap();
// Image 1 was examined and holds a face; 2 was examined and found
// empty; 3 holds a face found by a different model.
for (image, model) in [(1, "m"), (2, "m"), (3, "m")] {
c.execute(
"INSERT INTO face_index(image_id, model_id, indexed_at, faces_found, source_edge)
VALUES (?1, ?2, 0, 0, 2560)",
rusqlite::params![image, model],
)
.unwrap();
}
for (image, model) in [(1, "m"), (3, "other")] {
c.execute(
"INSERT INTO faces
(image_id, x, y, w, h, landmarks, detector_confidence, embedding,
crop_px, model_id, detected_at)
VALUES (?1, 0.1, 0.1, 0.2, 0.2, X'00', 0.9, X'00', 180.0, ?2, 0)",
rusqlite::params![image, model],
)
.unwrap();
}
c.pragma_update(None, "user_version", 13).unwrap();
migrate(&c).unwrap();
let kept: Vec<i64> = c
.prepare("SELECT image_id FROM face_index ORDER BY image_id")
.unwrap()
.query_map([], |r| r.get(0))
.unwrap()
.map(Result::unwrap)
.collect();
// 1 goes: it has a face with no quality. 2 stays: nothing on it to
// measure. 3 stays: its face belongs to a run this marker does not
// describe.
assert_eq!(kept, vec![2, 3]);
// And the faces themselves are untouched.
let faces: i64 = c
.query_row("SELECT count(*) FROM faces", [], |r| r.get(0))
.unwrap();
assert_eq!(faces, 2);
}
#[test]
fn job_uniqueness_coalesces_rather_than_duplicating() {
let c = mem();
+22
View File
@@ -54,9 +54,31 @@ pub fn checkpoint(conn: &Connection) -> Result<(), CatalogError> {
pub fn snapshot_for_upload(conn: &Connection, dest: &Path) -> Result<(), CatalogError> {
let out = copy_to(conn, dest)?;
strip_face_crops(&out)?;
verify_snapshot(&out)?;
Ok(())
}
/// TRACES: NFR-R2
/// Refuse to hand over a snapshot that will not pass `quick_check`.
///
/// The upload is the copy every other device merges from, and a damaged one
/// costs far more than the check: each device downloads it, fails, and — for
/// a week, once — declines to push over it. `quick_check` reads every page
/// but skips index verification, which is the affordable version of "is this
/// a database" on a 40 MB file that has just been written and is still in the
/// page cache. A failure here is [`CatalogError::Corrupt`], the same thing a
/// receiving device would have said, so the sync reports it the same way.
fn verify_snapshot(snapshot: &Connection) -> Result<(), CatalogError> {
let verdict: String = snapshot.query_row("PRAGMA quick_check", [], |r| r.get(0))?;
if verdict == "ok" {
Ok(())
} else {
Err(CatalogError::Corrupt {
detail: format!("the snapshot for upload failed quick_check: {verdict}"),
})
}
}
/// Checkpoint, then copy the whole database to `dest`, and hand back the
/// connection to the copy.
///
+221
View File
@@ -0,0 +1,221 @@
//! TRACES: S15 | FR-MRG-3
//! Spike S15.1 — does rawler read back a linear DNG this application writes?
//!
//! cargo run -p dr-decode --example linear_dng [-- <out.dng>]
//!
//! Decides FR-MRG-3's container. A panorama composite is three linear samples
//! per pixel with a camera matrix attached, which is exactly what a
//! `LinearRaw` DNG is; if rawler parses one, the composite re-enters the
//! library as `Format::Dng` and the only new decode work is a `cpp == 3`
//! branch. If it does not, the container is a float TIFF with a decode path
//! of its own.
//!
//! The file is hand-rolled rather than written with the `tiff` crate, whose
//! encoder fixes `PhotometricInterpretation` to RGB and cannot say
//! `LinearRaw`. Eighty lines of IFD is the cheaper thing to own than a fork.
use rawler::rawsource::RawSource;
const W: u32 = 64;
const H: u32 = 48;
fn main() {
let bytes = write_linear_dng(W, H);
if let Some(path) = std::env::args().nth(1) {
std::fs::write(&path, &bytes).expect("write");
println!("wrote {path} ({} bytes)", bytes.len());
}
let source = RawSource::new_from_slice(&bytes);
let decoder = match rawler::get_decoder(&source) {
Ok(d) => d,
Err(e) => {
println!("FAIL get_decoder: {e}");
std::process::exit(1);
}
};
println!("ok decoder found");
let image = match decoder.raw_image(&source, &Default::default(), false) {
Ok(i) => i,
Err(e) => {
println!("FAIL raw_image: {e}");
std::process::exit(1);
}
};
println!(
"ok raw_image: {}×{}, cpp {}, bps {}, {} samples, make {:?} model {:?}",
image.width,
image.height,
image.cpp,
image.bps,
match &image.data {
rawler::RawImageData::Integer(v) => v.len(),
rawler::RawImageData::Float(v) => v.len(),
},
image.make,
image.model
);
println!(
" white {:?} black {:?} wb {:?}",
image.whitelevel.0,
image
.blacklevel
.levels
.iter()
.map(|r| r.n as f32 / r.d.max(1) as f32)
.collect::<Vec<_>>(),
image.wb_coeffs
);
// The pixel at (1, 0) was written as (1000, 2000, 3000): if the samples
// come back interleaved in that order, cpp == 3 means what it says.
if let rawler::RawImageData::Integer(v) = &image.data {
let i = image.cpp;
println!(" pixel (1,0) = {:?}", &v[i..i + image.cpp.min(3)]);
}
// What dr-decode itself makes of it: the colour matrix rawler parsed into
// the camera definition, and the profile the decoder would build from it.
println!(" rawler color_matrix: {:?}", image.camera.color_matrix);
let dng = dr_decode::profile::read_dng_matrices(decoder.as_ref());
let profile = dr_decode::CameraProfile::extract(&image, &dng);
println!(
" CameraProfile: {}",
profile
.as_ref()
.map(|p| format!("xyz_to_cam {:?}", p.xyz_to_cam()))
.unwrap_or_else(|| "none".into())
);
match dr_decode::decode(&bytes) {
Ok(r) => println!(
"note dr_decode::decode accepted it as CFA: {}×{}, {} samples — the cpp==3 branch is the work",
r.width,
r.height,
r.data.len()
),
Err(e) => println!("note dr_decode::decode refused it: {e} — the cpp==3 branch is the work"),
}
}
/// A minimal `LinearRaw` DNG: one IFD, uncompressed 16-bit RGB, the tags a
/// decoder needs to treat it as a DNG and the matrix a develop chain needs
/// to treat it as a camera. Little-endian, one strip.
fn write_linear_dng(w: u32, h: u32) -> Vec<u8> {
// Pixels first, so their offset is known: a ramp with one marker pixel.
let mut pixels: Vec<u16> = Vec::with_capacity((w * h * 3) as usize);
for y in 0..h {
for x in 0..w {
if (x, y) == (1, 0) {
pixels.extend([1000, 2000, 3000]);
} else {
let v = ((x + y) * 512).min(65535) as u16;
pixels.extend([v, v / 2, v / 3]);
}
}
}
let pixel_bytes: Vec<u8> = pixels.iter().flat_map(|v| v.to_le_bytes()).collect();
// Layout: header (8) | pixels | extra data | IFD.
let pixels_off = 8u32;
let extra_off = pixels_off + pixel_bytes.len() as u32;
// Values that do not fit in four bytes go in `extra`, and the entry
// points at them.
let mut extra: Vec<u8> = Vec::new();
let mut entries: Vec<(u16, u16, u32, [u8; 4])> = Vec::new();
fn short(tag: u16, v: u16) -> (u16, u16, u32, [u8; 4]) {
let mut b = [0u8; 4];
b[..2].copy_from_slice(&v.to_le_bytes());
(tag, 3, 1, b)
}
fn long(tag: u16, v: u32) -> (u16, u16, u32, [u8; 4]) {
(tag, 4, 1, v.to_le_bytes())
}
fn ascii(extra: &mut Vec<u8>, extra_off: u32, tag: u16, s: &str) -> (u16, u16, u32, [u8; 4]) {
let mut bytes = s.as_bytes().to_vec();
bytes.push(0);
let off = extra_off + extra.len() as u32;
extra.extend(&bytes);
(tag, 2, bytes.len() as u32, off.to_le_bytes())
}
entries.push(long(254, 0)); // NewSubfileType: main image
entries.push(long(256, w));
entries.push(long(257, h));
// BitsPerSample ×3 — three shorts, six bytes, so out of line.
{
let off = extra_off + extra.len() as u32;
for _ in 0..3 {
extra.extend(16u16.to_le_bytes());
}
entries.push((258, 3, 3, off.to_le_bytes()));
}
entries.push(short(259, 1)); // Compression: none
entries.push(short(262, 34892)); // PhotometricInterpretation: LinearRaw
entries.push(ascii(&mut extra, extra_off, 271, "DarkRoom"));
entries.push(ascii(&mut extra, extra_off, 272, "Panorama"));
entries.push(long(273, pixels_off)); // StripOffsets
entries.push(short(274, 1)); // Orientation
entries.push(short(277, 3)); // SamplesPerPixel
entries.push(long(278, h)); // RowsPerStrip
entries.push(long(279, pixel_bytes.len() as u32)); // StripByteCounts
entries.push(short(284, 1)); // PlanarConfiguration: chunky
entries.push((50706, 1, 4, [1, 4, 0, 0])); // DNGVersion
entries.push((50707, 1, 4, [1, 4, 0, 0])); // DNGBackwardVersion
entries.push(ascii(&mut extra, extra_off, 50708, "DarkRoom Panorama")); // UniqueCameraModel
entries.push(long(50717, 65535)); // WhiteLevel
// ColorMatrix1: XYZ → camera, 9 SRATIONALs. A plausible sRGB-ish matrix
// (the inverse of the sRGB D65 primaries), scaled to integers.
{
let m: [(i32, i32); 9] = [
(32406, 10000),
(-15372, 10000),
(-4986, 10000),
(-9689, 10000),
(18758, 10000),
(415, 10000),
(557, 10000),
(-2040, 10000),
(10570, 10000),
];
let off = extra_off + extra.len() as u32;
for (n, d) in m {
extra.extend(n.to_le_bytes());
extra.extend(d.to_le_bytes());
}
entries.push((50721, 10, 9, off.to_le_bytes()));
}
// AsShotNeutral: 3 RATIONALs, neutral.
{
let off = extra_off + extra.len() as u32;
for _ in 0..3 {
extra.extend(1u32.to_le_bytes());
extra.extend(1u32.to_le_bytes());
}
entries.push((50728, 5, 3, off.to_le_bytes()));
}
entries.push(short(50778, 21)); // CalibrationIlluminant1: D65
entries.sort_by_key(|e| e.0);
let ifd_off = extra_off + extra.len() as u32;
let mut out = Vec::new();
out.extend(b"II");
out.extend(42u16.to_le_bytes());
out.extend(ifd_off.to_le_bytes());
out.extend(&pixel_bytes);
out.extend(&extra);
out.extend((entries.len() as u16).to_le_bytes());
for (tag, ty, count, value) in &entries {
out.extend(tag.to_le_bytes());
out.extend(ty.to_le_bytes());
out.extend(count.to_le_bytes());
out.extend(value);
}
out.extend(0u32.to_le_bytes()); // no next IFD
out
}
+60
View File
@@ -25,6 +25,42 @@ pub enum DecodeError {
CorruptPreview(String),
}
/// Run a decoder call, and return a panic inside it as an error.
///
/// TRACES: FR-RAW-4 | NFR-SEC-1 | NFR-R3
/// rawler `panic!`s on some malformed input rather than returning `Err` — a
/// DNG whose IFD claims a >50000 px image, for one, which is in the reference
/// library. A panic on a worker thread ends the thread: the face sweep that
/// met that file stopped 13 seconds in, three sweeps running, with "17301
/// image(s) to index" as the last word and nothing to say why. FR-RAW-4's
/// rule — a malformed file must not abort a batch — is this crate's to keep
/// whatever the library beneath it does, so every entry point that calls into
/// rawler runs through here, and a file that panics the decoder is one failed
/// file like any other.
///
/// The crash hook still records the panic, because it runs before unwinding
/// reaches this frame; that is right — it is a real defect in a dependency
/// and the record is how it gets reported upstream — and a repeat is the same
/// file being met again rather than a new fault.
pub(crate) fn guarded<T>(
what: &'static str,
f: impl FnOnce() -> Result<T, DecodeError>,
) -> Result<T, DecodeError> {
match std::panic::catch_unwind(std::panic::AssertUnwindSafe(f)) {
Ok(result) => result,
Err(payload) => {
let msg = payload
.downcast_ref::<&str>()
.map(|s| s.to_string())
.or_else(|| payload.downcast_ref::<String>().cloned())
.unwrap_or_else(|| "no message".to_string());
Err(DecodeError::Decode(format!(
"{what}: the decoder panicked on this file: {msg}"
)))
}
}
}
impl DecodeError {
/// Whether a fallback path might still produce an image.
///
@@ -49,4 +85,28 @@ mod tests {
// A genuinely unsupported file has nowhere to fall through to.
assert!(!DecodeError::Unsupported("unknown".into()).has_fallback());
}
#[test]
fn a_panic_in_the_decoder_is_an_error_and_the_thread_survives() {
// The property the face sweep relies on: one file that panics rawler
// is one failed file, not the end of the pass. The message travels,
// because "decode failed" alone sends the reader to the crash log.
let err = guarded("decode", || -> Result<(), DecodeError> {
panic!("rawler: surely there's no such thing as a {}MP image!", 600)
})
.unwrap_err();
let text = err.to_string();
assert!(text.contains("panicked"), "{text}");
assert!(text.contains("600MP"), "{text}");
assert!(!err.has_fallback(), "a panic is not a missing preview");
}
#[test]
fn a_result_passes_through_untouched() {
assert_eq!(guarded("decode", || Ok::<_, DecodeError>(7)).unwrap(), 7);
assert!(matches!(
guarded("decode", || Err::<(), _>(DecodeError::NoPreview)),
Err(DecodeError::NoPreview)
));
}
}
+43
View File
@@ -134,6 +134,22 @@ pub struct RawImage {
pub base_curve: BaseCurve,
/// The usable region of `data`, excluding masked and border photosites.
pub crop: CropRect,
/// TRACES: FR-MRG-3
/// Samples per photosite in `data`: 1 for a colour-filter-array capture,
/// 3 for a *linear* DNG — demosaiced RGB, still camera-space, which is
/// what a merge writes. With 3, `cfa_pattern` means nothing, `data` is
/// `width × height × 3` interleaved, and the GPU uploads it as it is
/// rather than demosaicing.
pub samples_per_pixel: u8,
/// TRACES: FR-MRG-3
/// The body's colour profile as the file carried it, for a composite to
/// carry on: calibrations and the as-shot neutral. `None` for a body the
/// decoder has no matrix for.
pub profile: Option<profile::CameraProfile>,
/// The body, as rawler cleans the names: what `Make`/`Model` say and what
/// the base-curve database matches on.
pub make: String,
pub model: String,
}
/// TRACES: FR-RAW-3
@@ -285,6 +301,10 @@ pub fn probe(header: &[u8]) -> Option<Format> {
/// TRACES: FR-CAT-5 | M-12
/// Read capture metadata without decoding sensor data.
pub fn metadata(bytes: &[u8]) -> Result<Metadata, DecodeError> {
error::guarded("metadata", || metadata_unguarded(bytes))
}
fn metadata_unguarded(bytes: &[u8]) -> Result<Metadata, DecodeError> {
use rawler::rawsource::RawSource;
// rawler has no decoder for a plain JPEG, so without this every JPEG in a
@@ -510,6 +530,10 @@ pub(crate) fn parse_exif_offset(s: &str) -> Option<i32> {
/// Only develop and export should call it; culling and the grid must not
/// (FR-CULL-1).
pub fn decode(bytes: &[u8]) -> Result<RawImage, DecodeError> {
error::guarded("decode", || decode_unguarded(bytes))
}
fn decode_unguarded(bytes: &[u8]) -> Result<RawImage, DecodeError> {
use rawler::rawsource::RawSource;
let source = RawSource::new_from_slice(bytes);
@@ -551,6 +575,21 @@ pub fn decode(bytes: &[u8]) -> Result<RawImage, DecodeError> {
image.camera.clean_model.as_str(),
);
// TRACES: FR-MRG-3
// A linear DNG — three samples per pixel, no colour filter array — is a
// composite this application wrote (or any other demosaiced DNG). It
// carries the same scale, matrices and neutral as a CFA file and goes
// through the same profile; only the demosaic is skipped.
let samples_per_pixel = match image.cpp {
1 => 1u8,
3 => 3,
other => {
return Err(DecodeError::Unsupported(format!(
"{other} samples per pixel; only CFA (1) and linear RGB (3) are handled"
)))
}
};
let data = match image.data {
rawler::RawImageData::Integer(v) => v,
rawler::RawImageData::Float(v) => {
@@ -619,6 +658,10 @@ pub fn decode(bytes: &[u8]) -> Result<RawImage, DecodeError> {
wb_coeffs,
color_matrix,
base_curve,
samples_per_pixel,
profile,
make: image.camera.clean_make.clone(),
model: image.camera.clean_model.clone(),
})
}
+4
View File
@@ -158,6 +158,10 @@ pub enum PreviewSize {
/// Returns [`DecodeError::NoPreview`] where there is none at all: a
/// fall-through signal, not a failure (see [`DecodeError::has_fallback`]).
pub fn extract_preview(bytes: &[u8], size: PreviewSize) -> Result<Preview, DecodeError> {
crate::error::guarded("preview", || extract_preview_unguarded(bytes, size))
}
fn extract_preview_unguarded(bytes: &[u8], size: PreviewSize) -> Result<Preview, DecodeError> {
use rawler::rawsource::RawSource;
// A plain JPEG *is* its own preview — rawler has no decoder for one, and
+48
View File
@@ -392,6 +392,21 @@ impl CameraProfile {
}
/// The calibrations this profile was built from, coolest first.
/// TRACES: FR-MRG-3
/// The calibrations as a DNG carries them: `(CalibrationIlluminant,
/// ColorMatrix)` with the EXIF light-source code, for a composite to
/// write the profile of the body that took its sources.
///
/// The code is recovered from the temperature, which is lossy only for
/// illuminants this profile never kept: `extract` drops calibrations
/// whose illuminant has no temperature, so every one here maps back.
pub fn dng_calibrations(&self) -> Vec<(u16, [[f32; 3]; 3])> {
self.calibrations
.iter()
.map(|c| (illuminant_code(c.temperature), c.xyz_to_cam))
.collect()
}
pub fn calibrations(&self) -> &[Calibration] {
&self.calibrations
}
@@ -528,6 +543,39 @@ fn illuminant_temperature(illuminant: Illuminant) -> Option<f32> {
})
}
/// The EXIF `LightSource` code for a calibration temperature — the inverse
/// of [`illuminant_temperature`], on the temperatures it produces.
fn illuminant_code(temperature: f32) -> u16 {
// Nearest of the table, so a temperature that came through a float
// round-trip still lands on its illuminant. Where two illuminants share
// a temperature (D55 and Daylight, D65 and Cloudy, D75 and Shade) the
// CIE standard one is written: it is what every profile database means.
const TABLE: &[(f32, u16)] = &[
(2856.0, 17), // A
(3200.0, 24), // ISO studio tungsten
(3500.0, 15), // white fluorescent
(4150.0, 14), // cool white fluorescent
(4230.0, 2), // fluorescent
(4874.0, 18), // B
(5000.0, 13), // daylight white fluorescent
(5003.0, 23), // D50
(5503.0, 20), // D55
(6430.0, 12), // daylight fluorescent
(6504.0, 21), // D65
(6774.0, 19), // C
(7504.0, 22), // D75
];
TABLE
.iter()
.min_by(|a, b| {
(a.0 - temperature)
.abs()
.total_cmp(&(b.0 - temperature).abs())
})
.map(|(_, code)| *code)
.unwrap_or(255)
}
/// Compose a forward matrix into camera RGB → linear sRGB.
///
/// `forward` takes white-balanced camera RGB to XYZ under D50, which is the
+3
View File
@@ -38,4 +38,7 @@ dr-gpu.workspace = true
dr-pipeline.workspace = true
env_logger.workspace = true
pollster.workspace = true
# The DNG writer's test reads its output back through the decoder the
# library uses, which is the whole claim the writer makes (S15.1).
rawler.workspace = true
zune-jpeg.workspace = true
+323
View File
@@ -0,0 +1,323 @@
//! TRACES: FR-MRG-3
//! A linear DNG: the container a merge writes its composite into.
//!
//! Decided by S15.1 (2026-09-19): rawler reads back a `LinearRaw` DNG the
//! application writes, so a composite re-enters the library as
//! `Format::Dng` through the decoder every camera DNG uses. What is written
//! is a RAW in every sense a warp can preserve — camera-linear `u16`
//! samples at the first source's own scale, its matrices, illuminants,
//! as-shot neutral and body name — so the panorama is developed afterwards
//! as one photograph, from the sensor's numbers.
//!
//! # Streamed, not buffered
//!
//! The composite is larger than any single photograph the pipeline renders
//! and larger than the tablet's memory (FR-MRG-11), so the writer never
//! holds it. Strips are pulled from the caller one at a time through a
//! closure, in order, and written as they arrive; the caller renders a band
//! of chunks, hands over its rows, and moves on.
//!
//! # Why the `tiff` crate after all
//!
//! S15.1's spike hand-rolled its IFD because the crate's encoder fixes
//! `PhotometricInterpretation` to RGB when the image is opened. It does — but
//! a directory is a map and a later `write_tag` on the same tag replaces the
//! earlier, so `LinearRaw` goes in over the top and everything else the
//! crate does (strips, offsets, sub-IFDs, the EXIF block `encode.rs` already
//! knows how to write) is kept.
use std::io::{Seek, Write};
use tiff::encoder::{colortype, DirectoryEncoder, SRational, TiffEncoder, TiffKind, TiffValue};
use tiff::tags::Tag;
use crate::encode::{sub_directories, tag_metadata, Ascii, Rationals};
use crate::{ExportError, SourceMetadata};
/// What the DNG says about the camera that "took" the composite: the first
/// source's profile, carried across so the composite develops through it.
#[derive(Debug, Clone, PartialEq)]
pub struct DngProfile {
/// `UniqueCameraModel`, the name the profile database matches on.
pub unique_model: String,
/// `(CalibrationIlluminant, ColorMatrix)`: the EXIF light-source code and
/// the XYZ → camera matrix measured under it. One or two.
pub calibrations: Vec<(u16, [[f32; 3]; 3])>,
/// `AsShotNeutral`, camera RGB of the scene's white.
pub as_shot_neutral: [f32; 3],
/// `WhiteLevel`: the sample value that is clipping. The first source's
/// white minus its black, since the samples are black-subtracted.
pub white_level: u32,
}
/// Write a linear DNG, pulling `rows_per_strip`-row strips from `strips`.
///
/// Each call to `strips` receives the strip index and a buffer to fill with
/// `width × rows × 3` interleaved RGB `u16` samples (the last strip may be
/// shorter). `source` supplies the `Make`, `Model`, dates and EXIF block
/// exactly as an export does (FR-EXP-8 sanitising already applied by the
/// caller).
///
/// `PhotometricInterpretation = LinearRaw`, `DNGVersion 1.4`, uncompressed,
/// `Orientation = 1` — the composite is written upright (panorama.md §8).
///
/// `crop` is asked once every strip is in, and its answer — the largest
/// rectangle the frames covered, found while the strips went by
/// (`Inscribed`) — becomes `DefaultCropOrigin`/`DefaultCropSize`
/// (FR-MRG-4): the file opens on the picture, and the border is still in it.
// Eight arguments, and each is a different thing: the sink, three
// dimensions, the profile, the header, the strip source and the crop. A
// struct for them would be a struct with one caller.
#[allow(clippy::too_many_arguments)]
pub fn write_linear_dng<W, F, C>(
out: W,
width: u32,
height: u32,
rows_per_strip: u32,
profile: &DngProfile,
source: Option<&SourceMetadata>,
mut strips: F,
crop: C,
) -> Result<(), ExportError>
where
W: Write + Seek,
F: FnMut(usize, &mut Vec<u16>) -> Result<(), ExportError>,
C: FnOnce() -> Option<crate::Rect>,
{
let enc = |e: tiff::TiffError| ExportError::Encode(e.to_string());
let mut encoder = TiffEncoder::new(out).map_err(enc)?;
let sub = sub_directories(&mut encoder, source, width, height)?;
let mut image = encoder
.new_image::<colortype::RGB16>(width, height)
.map_err(enc)?;
image.rows_per_strip(rows_per_strip.max(1)).map_err(enc)?;
tag_metadata(image.encoder(), source, &sub)?;
tag_dng(image.encoder(), profile).map_err(enc)?;
let rows = rows_per_strip.max(1);
let strip_count = height.div_ceil(rows) as usize;
let mut buf: Vec<u16> = Vec::with_capacity((width * rows * 3) as usize);
for k in 0..strip_count {
buf.clear();
strips(k, &mut buf)?;
let expected_rows = rows.min(height - k as u32 * rows);
let expected = (width * expected_rows * 3) as usize;
if buf.len() != expected {
return Err(ExportError::Encode(format!(
"strip {k} has {} samples, expected {expected}",
buf.len()
)));
}
image.write_strip(&buf).map_err(enc)?;
}
if let Some(r) = crop().filter(|r| r.width > 0 && r.height > 0) {
let r = crate::Rect {
x: r.x.min(width - 1),
y: r.y.min(height - 1),
width: r.width.min(width - r.x.min(width - 1)),
height: r.height.min(height - r.y.min(height - 1)),
};
image
.encoder()
.write_tag(Tag::Unknown(tag::DEFAULT_CROP_ORIGIN), &[r.x, r.y][..])
.map_err(enc)?;
image
.encoder()
.write_tag(
Tag::Unknown(tag::DEFAULT_CROP_SIZE),
&[r.width, r.height][..],
)
.map_err(enc)?;
}
image.finish().map_err(enc)
}
/// The tags that make a TIFF a DNG, and a linear one.
fn tag_dng<W, K>(dir: &mut DirectoryEncoder<'_, W, K>, profile: &DngProfile) -> tiff::TiffResult<()>
where
W: Write + Seek,
K: TiffKind,
{
// Over the top of what `new_image` wrote: this is the whole trick.
dir.write_tag(Tag::PhotometricInterpretation, LINEAR_RAW)?;
dir.write_tag(Tag::Orientation, 1u16)?;
dir.write_tag(Tag::Unknown(tag::DNG_VERSION), &[1u8, 4, 0, 0][..])?;
dir.write_tag(Tag::Unknown(tag::DNG_BACKWARD_VERSION), &[1u8, 4, 0, 0][..])?;
dir.write_tag(
Tag::Unknown(tag::UNIQUE_CAMERA_MODEL),
Ascii(&profile.unique_model),
)?;
dir.write_tag(
Tag::Unknown(tag::WHITE_LEVEL),
&[profile.white_level; 3][..],
)?;
dir.write_tag(Tag::Unknown(tag::BLACK_LEVEL), &[0u32; 3][..])?;
for (slot, (illuminant, matrix)) in profile.calibrations.iter().take(2).enumerate() {
let (ill_tag, mat_tag) = if slot == 0 {
(tag::CALIBRATION_ILLUMINANT_1, tag::COLOR_MATRIX_1)
} else {
(tag::CALIBRATION_ILLUMINANT_2, tag::COLOR_MATRIX_2)
};
dir.write_tag(Tag::Unknown(ill_tag), *illuminant)?;
let flat: Vec<SRational> = matrix
.iter()
.flatten()
.map(|&v| SRational {
n: (v * 10_000.0).round() as i32,
d: 10_000,
})
.collect();
dir.write_tag(Tag::Unknown(mat_tag), SRationals(&flat))?;
}
let neutral: Vec<(u32, u32)> = profile
.as_shot_neutral
.iter()
.map(|&v| ((v.max(0.0) * 1_000_000.0).round() as u32, 1_000_000))
.collect();
dir.write_tag(Tag::Unknown(tag::AS_SHOT_NEUTRAL), Rationals(&neutral))?;
Ok(())
}
/// `PhotometricInterpretation` for demosaiced, un-rendered sensor data.
const LINEAR_RAW: u16 = 34892;
/// DNG tag numbers the `tiff` crate has no names for.
mod tag {
pub const DNG_VERSION: u16 = 50706;
pub const DNG_BACKWARD_VERSION: u16 = 50707;
pub const UNIQUE_CAMERA_MODEL: u16 = 50708;
pub const BLACK_LEVEL: u16 = 50714;
pub const WHITE_LEVEL: u16 = 50717;
pub const DEFAULT_CROP_ORIGIN: u16 = 50719;
pub const DEFAULT_CROP_SIZE: u16 = 50720;
pub const COLOR_MATRIX_1: u16 = 50721;
pub const COLOR_MATRIX_2: u16 = 50722;
pub const AS_SHOT_NEUTRAL: u16 = 50728;
pub const CALIBRATION_ILLUMINANT_1: u16 = 50778;
pub const CALIBRATION_ILLUMINANT_2: u16 = 50779;
}
/// A run of `SRATIONAL`s, as `encode::Rationals` is for `RATIONAL`.
struct SRationals<'a>(&'a [SRational]);
impl TiffValue for SRationals<'_> {
const BYTE_LEN: u8 = 8;
const FIELD_TYPE: tiff::tags::Type = tiff::tags::Type::SRATIONAL;
fn count(&self) -> usize {
self.0.len()
}
fn data(&self) -> std::borrow::Cow<'_, [u8]> {
let mut out = Vec::with_capacity(self.0.len() * 8);
for r in self.0 {
out.extend_from_slice(&r.n.to_ne_bytes());
out.extend_from_slice(&r.d.to_ne_bytes());
}
std::borrow::Cow::Owned(out)
}
}
#[cfg(test)]
mod tests {
use super::*;
fn profile() -> DngProfile {
DngProfile {
unique_model: "Canon EOS 6D".into(),
calibrations: vec![
(17, [[0.8, -0.2, 0.1], [-0.3, 1.1, 0.2], [0.0, -0.1, 0.9]]),
(21, [[0.7, -0.1, 0.0], [-0.2, 1.0, 0.1], [0.0, -0.2, 0.8]]),
],
as_shot_neutral: [0.5, 1.0, 0.6],
white_level: 13_023,
}
}
fn write(width: u32, height: u32, rows: u32) -> Vec<u8> {
let mut bytes = std::io::Cursor::new(Vec::new());
let source = SourceMetadata {
make: Some("Canon".into()),
model: Some("Canon EOS 6D".into()),
..Default::default()
};
write_linear_dng(
&mut bytes,
width,
height,
rows,
&profile(),
Some(&source),
|k, buf| {
let first = k as u32 * rows;
let n = rows.min(height - first);
for y in first..first + n {
for x in 0..width {
buf.extend([(x + y * width) as u16, 1000, 2000]);
}
}
Ok(())
},
|| {
Some(crate::Rect {
x: 2,
y: 1,
width: 15,
height: 10,
})
},
)
.expect("written");
bytes.into_inner()
}
#[test]
fn rawler_reads_it_back_as_linear_raw() {
let bytes = write(20, 13, 4);
let source = rawler::rawsource::RawSource::new_from_slice(&bytes);
let decoder = rawler::get_decoder(&source).expect("a DNG");
let image = decoder
.raw_image(&source, &Default::default(), false)
.expect("decodes");
assert_eq!((image.width, image.height, image.cpp), (20, 13, 3));
assert_eq!(image.whitelevel.0[0], 13_023);
// Pixel (3, 2) is (3 + 2·20, 1000, 2000) — samples in order, strips
// joined without a seam.
let rawler::RawImageData::Integer(data) = &image.data else {
panic!("integer samples")
};
let i = (2 * 20 + 3) * 3;
assert_eq!(&data[i..i + 3], &[43, 1000, 2000]);
// Last row, from the short final strip.
let i = (12 * 20 + 19) * 3;
assert_eq!(data[i], (19 + 12 * 20) as u16);
// The profile came through as the camera's.
assert!(!image.camera.color_matrix.is_empty());
assert_eq!(image.model, "Canon EOS 6D");
// The default crop is what the decoder reports as the picture.
let crop = image.crop_area.expect("a crop");
assert_eq!((crop.p.x, crop.p.y, crop.d.w, crop.d.h), (2, 1, 15, 10));
}
#[test]
fn a_strip_of_the_wrong_length_is_refused() {
let mut bytes = std::io::Cursor::new(Vec::new());
let err = write_linear_dng(
&mut bytes,
8,
8,
8,
&profile(),
None,
|_, buf| {
buf.extend([0u16; 10]);
Ok(())
},
|| None,
)
.unwrap_err();
assert!(matches!(err, ExportError::Encode(_)));
}
}
+5 -5
View File
@@ -213,7 +213,7 @@ impl tiff::encoder::TiffValue for Undefined<'_> {
/// specification says, `dr-decode` reads them back with `from_utf8_lossy`, and
/// a mangled accent is a far better outcome than a refusal. So the bytes go
/// through verbatim with the terminating NUL the type requires.
struct Ascii<'a>(&'a str);
pub(crate) struct Ascii<'a>(pub(crate) &'a str);
impl tiff::encoder::TiffValue for Ascii<'_> {
const BYTE_LEN: u8 = 1;
@@ -241,7 +241,7 @@ impl tiff::encoder::TiffValue for Ascii<'_> {
/// a value that forced little-endian would be read back byte-swapped on a
/// big-endian machine. `exif.rs` builds its own header and so chooses its own
/// order; here the container has already chosen.
struct Rationals<'a>(&'a [(u32, u32)]);
pub(crate) struct Rationals<'a>(pub(crate) &'a [(u32, u32)]);
impl tiff::encoder::TiffValue for Rationals<'_> {
const BYTE_LEN: u8 = 8;
@@ -317,7 +317,7 @@ where
/// and then no pointer is written either, so the file has no trace of the
/// directory rather than a pointer to an empty one.
#[derive(Default)]
struct SubDirectories {
pub(crate) struct SubDirectories {
exif: Option<u32>,
gps: Option<u32>,
}
@@ -335,7 +335,7 @@ struct SubDirectories {
/// A TIFF gets no separate EXIF *block* — no APP1, no `eXIf` chunk. Its own
/// directory is the EXIF structure, and adding a second copy inside it would
/// give a reader two answers to every question.
fn sub_directories<W>(
pub(crate) fn sub_directories<W>(
encoder: &mut tiff::encoder::TiffEncoder<W>,
source: Option<&SourceMetadata>,
width: u32,
@@ -465,7 +465,7 @@ where
///
/// No `Orientation`, for the reason `exif.rs` gives at length: the pixels
/// arriving here are already upright.
fn tag_metadata<W, K>(
pub(crate) fn tag_metadata<W, K>(
dir: &mut tiff::encoder::DirectoryEncoder<'_, W, K>,
source: Option<&SourceMetadata>,
sub: &SubDirectories,
+156
View File
@@ -0,0 +1,156 @@
//! TRACES: FR-MRG-4
//! The largest rectangle inside a coverage mask, found a row at a time.
//!
//! A merged panorama has ragged edges: the frames' footprints under a
//! cylinder or a sphere are not rectangles, and the composite carries a
//! black border where none of them reached. FR-MRG-4 asks for an auto-crop
//! to the largest inscribed rectangle. This finds it as the bands are
//! produced, so the composite is never held to be measured (FR-MRG-11):
//! each row extends a running histogram of consecutive covered rows above
//! it, and the largest rectangle ending on that row is the largest
//! rectangle under the histogram — a stack pass, linear in the width.
//!
//! The crop is written as the DNG's `DefaultCropOrigin`/`DefaultCropSize`,
//! which every reader honours and which discards nothing: the pixels
//! outside it are still in the file for a photographer who wants them.
/// The rectangle so far, in pixels from the top left.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
pub struct Rect {
pub x: u32,
pub y: u32,
pub width: u32,
pub height: u32,
}
impl Rect {
pub fn area(&self) -> u64 {
u64::from(self.width) * u64::from(self.height)
}
}
/// Feed rows top to bottom; ask for the best at any point.
#[derive(Debug, Clone)]
pub struct Inscribed {
width: usize,
/// How many consecutive covered rows end at the last row fed, per column.
heights: Vec<u32>,
rows: u32,
best: Rect,
}
impl Inscribed {
pub fn new(width: u32) -> Self {
Inscribed {
width: width as usize,
heights: vec![0; width as usize],
rows: 0,
best: Rect::default(),
}
}
/// One more row of coverage, `width` long.
pub fn push_row(&mut self, covered: &[bool]) {
debug_assert_eq!(covered.len(), self.width);
for (h, &c) in self.heights.iter_mut().zip(covered) {
*h = if c { *h + 1 } else { 0 };
}
self.rows += 1;
// Largest rectangle under the histogram, with a sentinel column of
// height 0 at the end so every bar is popped.
let mut stack: Vec<usize> = Vec::new();
for i in 0..=self.width {
let h = if i < self.width { self.heights[i] } else { 0 };
while let Some(&top) = stack.last() {
if self.heights[top] <= h {
break;
}
stack.pop();
let height = self.heights[top];
let left = stack.last().map_or(0, |&l| l + 1);
let width = (i - left) as u32;
let area = u64::from(width) * u64::from(height);
if area > self.best.area() {
self.best = Rect {
x: left as u32,
y: self.rows - height,
width,
height,
};
}
}
stack.push(i);
}
}
/// Several rows at once, as a band hands them over.
pub fn push_rows(&mut self, covered: &[bool], rows: u32) {
for r in 0..rows as usize {
self.push_row(&covered[r * self.width..(r + 1) * self.width]);
}
}
pub fn best(&self) -> Rect {
self.best
}
}
#[cfg(test)]
mod tests {
use super::*;
fn from_art(art: &[&str]) -> Rect {
let mut ins = Inscribed::new(art[0].len() as u32);
for row in art {
let covered: Vec<bool> = row.chars().map(|c| c == '#').collect();
ins.push_row(&covered);
}
ins.best()
}
#[test]
fn a_full_mask_is_its_own_rectangle() {
let r = from_art(&["####", "####", "####"]);
assert_eq!(
r,
Rect {
x: 0,
y: 0,
width: 4,
height: 3
}
);
}
#[test]
fn ragged_edges_are_cut_off() {
// A cylinder's footprint: narrower at top and bottom.
let r = from_art(&[
"..####..", ".######.", "########", "########", ".######.", "..####..",
]);
// 6 wide × 4 tall = 24 beats 8 × 2 = 16 and 4 × 6 = 24 ties; the
// first found wins a tie, which is the wider one here.
assert_eq!(r.area(), 24);
assert!(r.width == 6 && r.height == 4 || r.width == 4 && r.height == 6);
}
#[test]
fn a_hole_is_avoided() {
let r = from_art(&["#####", "##.##", "#####", "#####"]);
// Left of the hole: 2 × 4 = 8; right: 2 × 4 = 8; below: 5 × 2 = 10.
assert_eq!(
r,
Rect {
x: 0,
y: 2,
width: 5,
height: 2
}
);
}
#[test]
fn nothing_covered_is_nothing() {
assert_eq!(from_art(&["....", "...."]).area(), 0);
}
}
+4
View File
@@ -24,16 +24,20 @@
use dr_types::{ColourSpace, ExportFormat, ExportSettings};
mod dng;
mod encode;
mod error;
mod exif;
pub mod icc;
mod inscribed;
mod metadata;
mod name;
mod sharpen;
mod size;
pub use dng::{write_linear_dng, DngProfile};
pub use error::ExportError;
pub use inscribed::{Inscribed, Rect};
pub use metadata::SourceMetadata;
pub use name::{resolve_name, NameContext};
pub use size::target_size;
+10 -5
View File
@@ -9,10 +9,11 @@ license.workspace = true
thiserror.workspace = true
log.workspace = true
# Inference. `ort` is the API; **tract is the engine** — see the workspace
# manifest, and docs/faces.md §3, for why the C++ ONNX Runtime is not linked.
# Inference. `ort` is the API; **what runs it is `dr-inference-engine`'s
# business** — tract, or an ONNX Runtime the app found on disk, on whichever
# provider the device has (docs/inference.md). This crate never names either.
ort = { workspace = true, optional = true }
ort-tract = { workspace = true, optional = true }
dr-inference-engine = { workspace = true, optional = true }
ndarray = { workspace = true, optional = true }
[dev-dependencies]
@@ -20,7 +21,7 @@ zune-jpeg.workspace = true
env_logger.workspace = true
# The M1 probe drives `ort` directly so it can print the raw load error.
ort = { workspace = true }
ort-tract = { workspace = true }
dr-inference-engine = { workspace = true }
[[example]]
name = "probe"
@@ -30,6 +31,10 @@ required-features = ["inference"]
name = "faces"
required-features = ["inference"]
[[example]]
name = "eyes"
required-features = ["inference"]
[features]
# Nothing on by default, and in particular **no `embedded-model`**: the weights
# are not a build input and never become one (docs/faces.md §2.2). A feature
@@ -44,4 +49,4 @@ default = []
# must be testable against synthetic embeddings on a machine with no weights on
# it — a test suite that needs a research-licensed download is a test suite
# that does not run in CI.
inference = ["dep:ort", "dep:ort-tract", "dep:ndarray"]
inference = ["dep:ort", "dep:dr-inference-engine", "dep:ndarray"]
+145
View File
@@ -0,0 +1,145 @@
//! Detect the faces in a JPEG and read each one's eyes (docs/faces.md §17).
//!
//! The thing worth looking at is whether the eye boxes land on eyes and
//! whether soft ones are refused — so with `--dump DIR` the crops the
//! classifiers were shown are written out as PPMs, one per eye and one per
//! head framing, named by image and face, and every line carries the
//! numbers the readability floors are set from.
//!
//! cargo run -p dr-face --features inference --example eyes -- \
//! DET.onnx 2D106DET.onnx OCEC.onnx SGC.onnx [--dump DIR] photo.jpg [photo.jpg ...]
//!
//! All four models must have had their dynamic dims pinned first; see
//! `tools/fix-face-model-shapes.sh`.
use std::path::{Path, PathBuf};
use std::time::Instant;
use dr_face::{align, DetectOptions, Detector, EyeModels, Pixels};
fn main() {
env_logger::init();
let mut args: Vec<String> = std::env::args().skip(1).collect();
let dump = args.iter().position(|a| a == "--dump").map(|i| {
args.remove(i);
PathBuf::from(args.remove(i))
});
if args.len() < 5 {
eprintln!(
"usage: eyes DET.onnx 2D106DET.onnx OCEC.onnx SGC.onnx [--dump DIR] IMAGE.jpg [IMAGE.jpg ...]"
);
std::process::exit(2);
}
if let Some(d) = &dump {
std::fs::create_dir_all(d).expect("dump dir");
}
let t = Instant::now();
let mut detector = Detector::from_path(&args[0]).expect("load detector");
let mut models = EyeModels::from_paths(&args[1], &args[2], &args[3]).expect("load eye models");
println!("loaded the models in {:?}", t.elapsed());
let opts = DetectOptions::default();
for path in &args[4..] {
let (rgb, w, h) = match load_jpeg(path) {
Ok(v) => v,
Err(e) => {
println!("{path}: {e}");
continue;
}
};
let dets = detector.detect(&rgb, w, h, &opts).expect("detect");
println!("\n{path} ({w}×{h}) {} face(s)", dets.len());
let stem = Path::new(path)
.file_stem()
.map(|s| s.to_string_lossy().into_owned())
.unwrap_or_default();
for (i, d) in dets.iter().enumerate() {
let px = Pixels::RgbF32(&rgb);
let t = Instant::now();
let reading = models
.read(px, w, h, d.bbox, &d.landmarks)
.expect("read eyes");
let ms = t.elapsed().as_secs_f64() * 1e3;
let Some((r, lm)) = reading else {
println!(" [{i}] nothing to cut, skipped");
continue;
};
println!(
" [{i}] conf {:.2} box {:.0}×{:.0} right {:.3} ({:.0}px, sharp {:.3}) left {:.3} ({:.0}px, sharp {:.3}) sunglasses {:.3} → {:?} ({ms:.1} ms)",
d.confidence,
d.width(),
d.height(),
r.right.open,
r.right.px,
r.right.sharpness,
r.left.open,
r.left.px,
r.left.sharpness,
r.sunglasses,
r.state(),
);
if let Some(dir) = &dump {
// The same crops `EyeModels::read` cut, cut again for the
// sheet from the landmarks it handed back: the reading itself
// carries numbers, not pixels.
for (name, contour) in [("right", lm.right_eye()), ("left", lm.left_eye())] {
if let Some(patch) =
align::eye_box(&contour).and_then(|b| align::eye_patch(px, w, h, b))
{
write_ppm(
&dir.join(format!("{stem}-{i}-{name}.ppm")),
patch.pixels(),
align::EYE_PATCH_WIDTH,
align::EYE_PATCH_HEIGHT,
);
}
}
if let Some(head) = align::head_views(px, w, h, &d.landmarks) {
for (n, view) in head.views().enumerate() {
write_ppm(
&dir.join(format!("{stem}-{i}-head{n}.ppm")),
view,
align::SUNGLASSES_EDGE,
align::SUNGLASSES_EDGE,
);
}
}
}
}
}
}
fn write_ppm(path: &Path, rgb: &[f32], w: usize, h: usize) {
let mut out = format!("P6\n{w} {h}\n255\n").into_bytes();
out.extend(
rgb.iter()
.map(|v| (v.clamp(0.0, 1.0) * 255.0).round() as u8),
);
std::fs::write(path, out).expect("write ppm");
}
/// Decode to the tightly packed `f32` RGB `0.0..=1.0` the crate expects.
fn load_jpeg(path: &str) -> Result<(Vec<f32>, usize, usize), String> {
let bytes = std::fs::read(path).map_err(|e| e.to_string())?;
let mut dec = zune_jpeg::JpegDecoder::new(&bytes);
let px = dec.decode().map_err(|e| e.to_string())?;
let info = dec.info().ok_or("no jpeg header")?;
let (w, h) = (info.width as usize, info.height as usize);
let rgb: Vec<f32> = match px.len() / (w * h) {
3 => px.iter().map(|&v| v as f32 / 255.0).collect(),
1 => px
.iter()
.flat_map(|&v| {
let g = v as f32 / 255.0;
[g, g, g]
})
.collect(),
n => return Err(format!("{n} components per pixel, expected 1 or 3")),
};
Ok((rgb, w, h))
}
+3 -2
View File
@@ -63,15 +63,16 @@ fn main() {
let embed_ms = t.elapsed().as_secs_f64() * 1e3;
println!(
" [{i}] conf {:.3} box {:.0},{:.0} {:.0}×{:.0} crop_px {:.0} embed {embed_ms:.0} ms",
" [{i}] conf {:.3} box {:.0},{:.0} {:.0}×{:.0} crop_px {:.0} quality {:.1} embed {embed_ms:.0} ms",
d.confidence,
d.bbox.0,
d.bbox.1,
d.width(),
d.height(),
aligned.source_px(),
emb.quality,
);
all.push((path.clone(), i, emb));
all.push((path.clone(), i, emb.embedding));
}
}
+2
View File
@@ -46,11 +46,13 @@ fn main() {
);
for n in sizes {
let (embeddings, crop_px, images) = population(n);
let gallery = vec![true; n];
let faces = Faces {
embeddings: &embeddings,
dim: EMBEDDING_DIM,
crop_px: &crop_px,
images: &images,
gallery: &gallery,
};
let start = std::time::Instant::now();
+561 -65
View File
@@ -117,51 +117,56 @@ impl Aligned112 {
/// `face_index --quality` prints the joint distribution so the two are
/// chosen together rather than each in ignorance of the other.
pub fn sharpness(&self) -> f32 {
let e = ALIGNED_EDGE;
let luma: Vec<f32> = self
.pixels
.chunks_exact(3)
.map(|p| 0.2126 * p[0] + 0.7152 * p[1] + 0.0722 * p[2])
.collect();
let (mut lap_sum, mut lap_sq) = (0.0_f64, 0.0_f64);
let (mut lum_sum, mut lum_sq) = (0.0_f64, 0.0_f64);
let mut n = 0.0_f64;
for y in 1..e - 1 {
for x in 1..e - 1 {
let i = y * e + x;
// Four-neighbour Laplacian. The 8-neighbour form is more
// sensitive to diagonal detail and also to noise, which on a
// high-ISO frame is exactly the thing that must not read as
// sharpness.
let lap = 4.0 * luma[i] - luma[i - 1] - luma[i + 1] - luma[i - e] - luma[i + e];
let lap = lap as f64;
lap_sum += lap;
lap_sq += lap * lap;
let l = luma[i] as f64;
lum_sum += l;
lum_sq += l * l;
n += 1.0;
}
}
if n == 0.0 {
return 0.0;
}
let lap_var = (lap_sq / n - (lap_sum / n).powi(2)).max(0.0);
let lum_var = (lum_sq / n - (lum_sum / n).powi(2)).max(0.0);
// A crop with no luma variation has no edges to find either, so the
// ratio is 0/0. Zero is the right answer: nothing there is a face.
if lum_var <= 1e-9 {
return 0.0;
}
(lap_var / lum_var) as f32
laplacian_ratio(&self.pixels, ALIGNED_EDGE, ALIGNED_EDGE)
}
}
/// Variance of the four-neighbour Laplacian over the variance of the luma,
/// for a `w × h` RGB crop — the measure [`Aligned112::sharpness`] describes,
/// shared with [`EyePatch::sharpness`].
fn laplacian_ratio(pixels: &[f32], w: usize, h: usize) -> f32 {
let luma: Vec<f32> = pixels
.chunks_exact(3)
.map(|p| 0.2126 * p[0] + 0.7152 * p[1] + 0.0722 * p[2])
.collect();
let (mut lap_sum, mut lap_sq) = (0.0_f64, 0.0_f64);
let (mut lum_sum, mut lum_sq) = (0.0_f64, 0.0_f64);
let mut n = 0.0_f64;
for y in 1..h.saturating_sub(1) {
for x in 1..w.saturating_sub(1) {
let i = y * w + x;
// Four-neighbour Laplacian. The 8-neighbour form is more
// sensitive to diagonal detail and also to noise, which on a
// high-ISO frame is exactly the thing that must not read as
// sharpness.
let lap = 4.0 * luma[i] - luma[i - 1] - luma[i + 1] - luma[i - w] - luma[i + w];
let lap = lap as f64;
lap_sum += lap;
lap_sq += lap * lap;
let l = luma[i] as f64;
lum_sum += l;
lum_sq += l * l;
n += 1.0;
}
}
if n == 0.0 {
return 0.0;
}
let lap_var = (lap_sq / n - (lap_sum / n).powi(2)).max(0.0);
let lum_var = (lum_sq / n - (lum_sum / n).powi(2)).max(0.0);
// A crop with no luma variation has no edges to find either, so the
// ratio is 0/0. Zero is the right answer: nothing there is a face.
if lum_var <= 1e-9 {
return 0.0;
}
(lap_var / lum_var) as f32
}
/// A similarity transform: rotation, uniform scale, translation.
///
/// Stored as the four independent parameters rather than a 2×3 matrix so that
@@ -285,34 +290,381 @@ pub fn warp(
height: usize,
landmarks: &[(f32, f32); 5],
) -> Option<Aligned112> {
if rgb.len() != width * height * 3 {
warp_pixels(Pixels::RgbF32(rgb), width, height, landmarks)
}
/// TRACES: FR-CULL-8
/// What the warp may sample, in whichever layout the caller already holds.
///
/// # Why the 8-bit variant exists
///
/// FR-CULL-8 requires the crop to come from the **native** render, and a native
/// render is large: a 24 MP frame is 96 MB as `RGBA8` and 288 MB converted to
/// the `f32` RGB this module was originally written against. Converting the
/// whole frame to sample 112×112 from it is three hundred megabytes allocated
/// to read about forty thousand pixels, per image, on a pass that runs over a
/// whole library — and on Android it is NFR-RES-2's budget spent outright.
///
/// So the warp reads whatever the caller has instead. It touches so few pixels
/// that the per-sample conversion is free, and the buffer never has to be
/// duplicated in another layout.
#[derive(Debug, Clone, Copy)]
pub enum Pixels<'a> {
/// Tightly packed `f32` RGB in `0.0..=1.0`, row-major.
RgbF32(&'a [f32]),
/// Tightly packed 8-bit RGBA, row-major. Alpha is ignored: a face crop has
/// no use for it and carrying it would change what the embedder receives.
Rgba8(&'a [u8]),
}
impl Pixels<'_> {
/// Whether the buffer is the size `width × height` implies.
fn fits(&self, width: usize, height: usize) -> bool {
match self {
Pixels::RgbF32(v) => v.len() == width * height * 3,
Pixels::Rgba8(v) => v.len() == width * height * 4,
}
}
/// One channel of one pixel, as `0.0..=1.0`. Outside the buffer reads black.
///
/// Public because the face *crop* stored for the People screen is cut from
/// the same buffer by the same caller, and it should not need a second
/// copy of this to do it.
pub fn channel(&self, w: usize, h: usize, x: isize, y: isize, c: usize) -> f32 {
if x < 0 || y < 0 || x >= w as isize || y >= h as isize {
return 0.0;
}
let i = y as usize * w + x as usize;
match self {
Pixels::RgbF32(v) => v[i * 3 + c],
Pixels::Rgba8(v) => v[i * 4 + c] as f32 / 255.0,
}
}
}
/// [`warp`], over any layout [`Pixels`] describes.
pub fn warp_pixels(
px: Pixels<'_>,
width: usize,
height: usize,
landmarks: &[(f32, f32); 5],
) -> Option<Aligned112> {
if !px.fits(width, height) {
return None;
}
let m = fit_similarity(landmarks, &ARCFACE_TEMPLATE)?;
let e = ALIGNED_EDGE;
let mut pixels = vec![0.0_f32; e * e * 3];
for v in 0..e {
for u in 0..e {
// Pixel centres, so the transform is not off by half a pixel —
// which is small enough to survive review and large enough to
// matter on a 40-pixel face.
let (x, y) = m.invert(u as f32 + 0.5, v as f32 + 0.5);
let (x, y) = (x - 0.5, y - 0.5);
let out = (v * e + u) * 3;
sample_bilinear(rgb, width, height, x, y, &mut pixels[out..out + 3]);
}
}
let window = TemplateWindow {
x: 0.0,
y: 0.0,
w: e as f32,
h: e as f32,
};
Some(Aligned112 {
pixels,
pixels: sample_window(px, width, height, &m, &window, e, e),
// The warp maps `scale` source pixels to one destination pixel, so the
// crop spans 112/scale of the source.
source_px: ALIGNED_EDGE as f32 / m.scale(),
})
}
fn sample_bilinear(rgb: &[f32], w: usize, h: usize, x: f32, y: f32, out: &mut [f32]) {
/// A rectangle in **template** coordinates — the 112-unit frame
/// [`ARCFACE_TEMPLATE`] is written in — that a crop is sampled from.
///
/// Every crop this module makes is one of these resampled through the same
/// fitted similarity: the aligned face is the window `(0, 0, 112, 112)`, an
/// eye is a small window around its template point, a head is a window larger
/// than the face. Stating them all in one frame is what lets a second crop be
/// added as a constant rather than a second warp, and what keeps them
/// consistent with each other — the eye window sits where the eye landmark
/// lands *after* alignment, so a tilted face gets an upright eye.
#[derive(Debug, Clone, Copy, PartialEq)]
struct TemplateWindow {
x: f32,
y: f32,
w: f32,
h: f32,
}
/// Resample `window` of the template frame into an `out_w × out_h` RGB buffer.
///
/// Bilinear, from the source, in one step — the property [`warp`] insists on,
/// and every crop through here inherits it. The output pixel `(u, v)` is placed
/// at its centre in the window, taken back through `m` to source coordinates,
/// and sampled there; the window's aspect is **not** preserved when it differs
/// from the output's, which is deliberate for the eye classifier (it was
/// trained on detector boxes resized the same way) and moot for the others.
fn sample_window(
px: Pixels<'_>,
width: usize,
height: usize,
m: &Similarity,
window: &TemplateWindow,
out_w: usize,
out_h: usize,
) -> Vec<f32> {
let mut pixels = vec![0.0_f32; out_w * out_h * 3];
let sx = window.w / out_w as f32;
let sy = window.h / out_h as f32;
for v in 0..out_h {
for u in 0..out_w {
// Pixel centres, so the transform is not off by half a pixel —
// which is small enough to survive review and large enough to
// matter on a 40-pixel face.
let tx = window.x + (u as f32 + 0.5) * sx;
let ty = window.y + (v as f32 + 0.5) * sy;
let (x, y) = m.invert(tx, ty);
let (x, y) = (x - 0.5, y - 0.5);
let out = (v * out_w + u) * 3;
sample_bilinear(px, width, height, x, y, &mut pixels[out..out + 3]);
}
}
pixels
}
// ── eyes ──────────────────────────────────────────────────────────────────
/// Width of an eye crop as the classifier reads it, in pixels. Fixed by the
/// OCEC input (`docs/faces.md` §17): 40 wide, 24 high.
pub const EYE_PATCH_WIDTH: usize = 40;
/// Height of an eye crop as the classifier reads it, in pixels.
pub const EYE_PATCH_HEIGHT: usize = 24;
/// How much an eye's box is grown beyond its lid contour, as a fraction of
/// its width and height on each side.
///
/// The classifier was trained on a whole-body detector's *eye* boxes — tight
/// round the palpebral fissure — and measured on 25 open-eyed faces from the
/// reference library, a tight box is what it wants: 22 of 25 read open at
/// 0 and 0.1, 18 at 0.4, 14 at 0.6 (docs/faces.md §17.2). A tenth, so a
/// contour landing a pixel short of the lashes still holds them.
pub const EYE_BOX_MARGIN: f32 = 0.1;
/// Height a shut eye's box is given, as a fraction of its width.
///
/// A closed eye's contour has no height. The box is given the height an
/// open eye of the same width would have, so the classifier sees the same
/// framing either way — which is what it was trained on.
pub const EYE_BOX_MIN_ASPECT: f32 = 0.4;
/// The box round an eye's lid contour, in the contour's own coordinates:
/// `(x, y, w, h)`.
///
/// Model-free: the contour is whatever the landmark model gave for the ten
/// (or so) points on the lids, in source pixels. `None` for an empty
/// contour or one with no width, which is what a hidden eye's collapsed
/// contour can come to.
pub fn eye_box(contour: &[(f32, f32)]) -> Option<(f32, f32, f32, f32)> {
let (mut x0, mut y0, mut x1, mut y1) = (f32::MAX, f32::MAX, f32::MIN, f32::MIN);
for &(x, y) in contour {
x0 = x0.min(x);
y0 = y0.min(y);
x1 = x1.max(x);
y1 = y1.max(y);
}
let w = x1 - x0;
if contour.is_empty() || w <= 0.0 || w.is_nan() {
return None;
}
let h = (y1 - y0).max(w * EYE_BOX_MIN_ASPECT);
let cy = (y0 + y1) / 2.0;
let (mx, my) = (w * EYE_BOX_MARGIN, h * EYE_BOX_MARGIN);
Some((x0 - mx, cy - h / 2.0 - my, w + 2.0 * mx, h + 2.0 * my))
}
/// One eye, resampled to the classifier's input.
///
/// Constructible only by [`eye_patch`], for the reason [`Aligned112`] is
/// only constructible by [`warp`]: the classifier accepting a plain buffer
/// would accept any 40×24 of anything, and its answer would still be a
/// plausible probability.
#[derive(Debug, Clone, PartialEq)]
pub struct EyePatch {
/// `24 × 40 × 3`, row-major RGB in `0.0..=1.0`.
pixels: Vec<f32>,
/// Source pixels across the box the patch was cut from.
source_px: f32,
}
impl EyePatch {
pub fn pixels(&self) -> &[f32] {
&self.pixels
}
/// Source pixels across the eye box — how much eye there was to read.
///
/// The classifier was trained down to eyes a dozen pixels wide, and
/// below that a crop is an interpolation of nothing; `crate::eyes` draws
/// the line. Zero when the box had no width, which is a hidden eye.
pub fn source_px(&self) -> f32 {
self.source_px
}
/// How sharp the eye the classifier is about to see actually is —
/// [`Aligned112::sharpness`]'s measure, over the patch.
///
/// The reason it exists is the reason the face's does: a soft eye is
/// not a closed one, but a classifier shown a smear says "closed" with
/// the same confidence it says anything, and the only defence is to
/// not ask. A face sharp enough to embed can still hold an eye too soft
/// to read — it is a fortieth of the face — so the measure is taken
/// here and not inherited from the crop.
pub fn sharpness(&self) -> f32 {
laplacian_ratio(&self.pixels, EYE_PATCH_WIDTH, EYE_PATCH_HEIGHT)
}
}
/// Cut an eye out of the source at the classifier's size, from an
/// axis-aligned box in source pixels — [`eye_box`]'s, as a rule.
///
/// Upright and from the frame, not through the face's alignment: the
/// classifier's training crops were detector boxes, and a landmark model's
/// contour already says where the eye is on a tilted head. Bilinear in one
/// step from the native buffer, so a large face gives real pixels; the
/// box's aspect is not preserved, which is what the training resize did.
pub fn eye_patch(
px: Pixels<'_>,
width: usize,
height: usize,
bbox: (f32, f32, f32, f32),
) -> Option<EyePatch> {
let pixels = crop_box(px, width, height, bbox, EYE_PATCH_WIDTH, EYE_PATCH_HEIGHT)?;
Some(EyePatch {
pixels,
source_px: bbox.2,
})
}
// ── sunglasses ────────────────────────────────────────────────────────────
/// Edge of the crop the sunglasses classifier reads. Fixed by the SGC input:
/// 48×48.
pub const SUNGLASSES_EDGE: usize = 48;
/// The windows read for the sunglasses classifier, in template units:
/// `(x, y, w, h)`.
///
/// **Two framings, and the classifier's answer is the higher of the two.**
/// It was trained on a whole-body detector's *head* boxes, and a head box
/// is not reproducible from five landmarks: how much hair and hat it took in
/// depended on the person. So it is shown the face twice — once as the
/// aligned crop itself, once shifted up and widened to take in hair and
/// hat at the cost of the chin, which is roughly where a head box falls —
/// and a pair of sunglasses counts if it looks like one in either.
///
/// Measured over 12 faces in sunglasses and 28 with plainly visible eyes
/// from the reference library (`examples/eyes.rs --head`), at the 0.5
/// threshold:
///
/// | window | sunglasses found | clear eyes kept |
/// |---|---|---|
/// | the aligned face, `(0, 0, 112, 112)` | 9 | 28 |
/// | a head, `(-5, -14, 122, 122)` | 6 | 27 |
/// | a larger head, `(-30, -55, 172, 190)` | 6 | 25 |
/// | **the higher of the first two** | **11** | 27 |
///
/// The face-tight crop alone was the best single framing, which was not the
/// expectation; the head framing found the sunglasses under a cap that the
/// face crop missed. The one clear-eyed face the pair loses wears a cap and
/// clear glasses, at 0.68. Erring towards "sunglasses" is the safe direction
/// for what this feeds: a face called sunglasses is left alone by the
/// eyes-open filter, where a pair of sunglasses missed hands the eye
/// classifier a lens to guess at (docs/faces.md §17).
pub const SUNGLASSES_WINDOWS: [(f32, f32, f32, f32); 2] =
[(0.0, 0.0, 112.0, 112.0), (-5.0, -14.0, 122.0, 122.0)];
/// The framings of one face the sunglasses classifier is shown.
///
/// A newtype for the reason [`EyePatch`] is one.
#[derive(Debug, Clone, PartialEq)]
pub struct HeadViews {
/// Each `48 × 48 × 3`, row-major RGB in `0.0..=1.0`.
views: Vec<Vec<f32>>,
}
impl HeadViews {
pub fn views(&self) -> impl Iterator<Item = &[f32]> {
self.views.iter().map(Vec::as_slice)
}
}
/// Cut the [`SUNGLASSES_WINDOWS`] out of the source, aligned, at the
/// classifier's size.
pub fn head_views(
px: Pixels<'_>,
width: usize,
height: usize,
landmarks: &[(f32, f32); 5],
) -> Option<HeadViews> {
head_views_in(px, width, height, landmarks, &SUNGLASSES_WINDOWS)
}
/// [`head_views`] over windows other than [`SUNGLASSES_WINDOWS`].
///
/// For measuring them, which is how the constant was chosen
/// (`examples/eyes.rs --head`); production callers use the constant.
pub fn head_views_in(
px: Pixels<'_>,
width: usize,
height: usize,
landmarks: &[(f32, f32); 5],
windows: &[(f32, f32, f32, f32)],
) -> Option<HeadViews> {
if !px.fits(width, height) || windows.is_empty() {
return None;
}
let m = fit_similarity(landmarks, &ARCFACE_TEMPLATE)?;
let views = windows
.iter()
.map(|&(x, y, w, h)| {
let window = TemplateWindow { x, y, w, h };
sample_window(
px,
width,
height,
&m,
&window,
SUNGLASSES_EDGE,
SUNGLASSES_EDGE,
)
})
.collect();
Some(HeadViews { views })
}
/// An axis-aligned crop of the source, resampled to `out_w × out_h` RGB.
///
/// `(x, y, w, h)` in source pixels; the aspect is not preserved when it
/// differs from the output's. Bilinear in one step, like every crop here;
/// pixels outside the source read black. What a landmark model trained on
/// detector boxes wants — upright, from the frame — as against the aligned
/// windows above.
pub fn crop_box(
px: Pixels<'_>,
width: usize,
height: usize,
(x, y, w, h): (f32, f32, f32, f32),
out_w: usize,
out_h: usize,
) -> Option<Vec<f32>> {
if !px.fits(width, height) || w <= 0.0 || h <= 0.0 {
return None;
}
let identity = Similarity {
a: 1.0,
b: 0.0,
tx: 0.0,
ty: 0.0,
};
let window = TemplateWindow { x, y, w, h };
Some(sample_window(
px, width, height, &identity, &window, out_w, out_h,
))
}
fn sample_bilinear(px: Pixels<'_>, w: usize, h: usize, x: f32, y: f32, out: &mut [f32]) {
let x0 = x.floor();
let y0 = y.floor();
let fx = x - x0;
@@ -321,13 +673,7 @@ fn sample_bilinear(rgb: &[f32], w: usize, h: usize, x: f32, y: f32, out: &mut [f
let y0 = y0 as isize;
for (c, o) in out.iter_mut().enumerate() {
let get = |xi: isize, yi: isize| -> f32 {
if xi < 0 || yi < 0 || xi >= w as isize || yi >= h as isize {
0.0
} else {
rgb[(yi as usize * w + xi as usize) * 3 + c]
}
};
let get = |xi: isize, yi: isize| -> f32 { px.channel(w, h, xi, yi, c) };
let top = get(x0, y0) * (1.0 - fx) + get(x0 + 1, y0) * fx;
let bot = get(x0, y0 + 1) * (1.0 - fx) + get(x0 + 1, y0 + 1) * fx;
*o = top * (1.0 - fy) + bot * fy;
@@ -435,6 +781,125 @@ mod tests {
}
}
/// A source whose red channel is its x coordinate and green its y, so a
/// crop's mean colour says where in the source it was taken from.
fn coordinate_image(w: usize, h: usize) -> Vec<f32> {
let mut rgb = vec![0.0_f32; w * h * 3];
for y in 0..h {
for x in 0..w {
rgb[(y * w + x) * 3] = x as f32 / w as f32;
rgb[(y * w + x) * 3 + 1] = y as f32 / h as f32;
}
}
rgb
}
fn mean_channel(px: &[f32], c: usize) -> f32 {
let n = px.len() / 3;
px.chunks_exact(3).map(|p| p[c]).sum::<f32>() / n as f32
}
/// The box is the contour's bounds, grown by the margin, and a shut
/// eye's flat contour is given an open eye's height.
#[test]
fn an_eye_box_holds_its_contour_with_a_margin() {
let open = [(100.0, 50.0), (110.0, 46.0), (120.0, 50.0), (110.0, 54.0)];
let (x, y, w, h) = eye_box(&open).unwrap();
assert!((w - 20.0 * (1.0 + 2.0 * EYE_BOX_MARGIN)).abs() < 1e-4);
assert!((h - 8.0 * (1.0 + 2.0 * EYE_BOX_MARGIN)).abs() < 1e-4);
assert!((x + w / 2.0 - 110.0).abs() < 1e-4);
assert!((y + h / 2.0 - 50.0).abs() < 1e-4);
let shut = [(100.0, 50.0), (110.0, 50.0), (120.0, 50.0)];
let (_, _, w2, h2) = eye_box(&shut).unwrap();
assert!((w2 - w).abs() < 1e-4, "same width");
assert!((h2 - 20.0 * EYE_BOX_MIN_ASPECT * (1.0 + 2.0 * EYE_BOX_MARGIN)).abs() < 1e-4);
assert!(eye_box(&[]).is_none());
assert!(eye_box(&[(5.0, 5.0), (5.0, 9.0)]).is_none(), "no width");
}
/// The patch is cut from the box it was given, upright, and knows how
/// many source pixels it spans.
#[test]
fn an_eye_patch_is_the_box_resampled() {
let (w, h) = (200, 200);
let rgb = coordinate_image(w, h);
let bbox = (60.0, 90.0, 30.0, 12.0);
let eye = eye_patch(Pixels::RgbF32(&rgb), w, h, bbox).unwrap();
assert_eq!(eye.pixels().len(), EYE_PATCH_WIDTH * EYE_PATCH_HEIGHT * 3);
assert_eq!(eye.source_px(), 30.0);
let cx = mean_channel(eye.pixels(), 0) * w as f32;
let cy = mean_channel(eye.pixels(), 1) * h as f32;
assert!((cx - 75.0).abs() < 0.6, "{cx}");
assert!((cy - 96.0).abs() < 0.6, "{cy}");
// No width, or a buffer that is not the size it claims: nothing.
assert!(eye_patch(Pixels::RgbF32(&rgb), w, h, (60.0, 90.0, 0.0, 12.0)).is_none());
assert!(eye_patch(Pixels::RgbF32(&rgb), 190, 200, bbox).is_none());
}
/// A soft eye scores lower than the same eye sharp, on the patch itself.
#[test]
fn an_eye_patchs_sharpness_falls_with_blur() {
let edge = 120;
let sharp = image(
edge,
|x, y| if (x / 5 + y / 5) % 2 == 0 { 0.9 } else { 0.1 },
);
let soft = blur(&blur(&sharp, edge), edge);
let bbox = (20.0, 40.0, 40.0, 24.0);
let a = eye_patch(Pixels::RgbF32(&sharp), edge, edge, bbox)
.unwrap()
.sharpness();
let b = eye_patch(Pixels::RgbF32(&soft), edge, edge, bbox)
.unwrap()
.sharpness();
assert!(a > b * 2.0, "sharp {a} should clearly beat blurred {b}");
}
/// The second sunglasses framing takes in more than the face — it starts
/// above the template's top edge and ends below its bottom — and the
/// first is the aligned face itself.
#[test]
fn the_head_views_are_the_face_and_a_wider_framing_of_it() {
let (w, h) = (300, 300);
let rgb = coordinate_image(w, h);
let lm = shifted_scaled(1.0, 100.0, 100.0, 0.0);
let head = head_views(Pixels::RgbF32(&rgb), w, h, &lm).unwrap();
let views: Vec<&[f32]> = head.views().collect();
let face = warp(&rgb, w, h, &lm).unwrap();
assert_eq!(views.len(), SUNGLASSES_WINDOWS.len());
for v in &views {
assert_eq!(v.len(), SUNGLASSES_EDGE * SUNGLASSES_EDGE * 3);
}
// The face view samples the same region as the aligned crop.
assert!((mean_channel(views[0], 0) - mean_channel(face.pixels(), 0)).abs() < 0.01);
assert!((mean_channel(views[0], 1) - mean_channel(face.pixels(), 1)).abs() < 0.01);
let (x, y, ww, hh) = SUNGLASSES_WINDOWS[1];
assert!(
x < 0.0 && y < 0.0,
"the window starts outside the face crop"
);
assert!(x + ww > ALIGNED_EDGE as f32, "and is wider than it");
assert!(y + hh < ALIGNED_EDGE as f32, "but stops short of the chin");
// Centred horizontally on the face, so the two share a mean x.
assert!((mean_channel(views[1], 0) - mean_channel(face.pixels(), 0)).abs() < 0.01);
// Its first row lies above the face's first row.
assert!(views[1][1] < face.pixels()[1]);
}
#[test]
fn degenerate_landmarks_yield_no_head_crop() {
let rgb = vec![0.5_f32; 64 * 64 * 3];
let degenerate = [(50.0, 50.0); 5];
assert!(head_views(Pixels::RgbF32(&rgb), 64, 64, &degenerate).is_none());
// And a buffer that is not the size it claims.
let lm = shifted_scaled(1.0, 0.0, 0.0, 0.0);
assert!(head_views(Pixels::RgbF32(&rgb), 60, 60, &lm).is_none());
}
#[test]
fn out_of_bounds_samples_read_black_rather_than_wrapping() {
let rgb = vec![1.0_f32; 32 * 32 * 3];
@@ -508,6 +973,37 @@ mod tests {
/// against a bright sky is low-contrast, and a raw Laplacian variance would
/// reject it as blurred — which would quietly throw away every backlit
/// portrait in the library.
#[test]
fn both_pixel_layouts_warp_to_the_same_crop() {
// The 8-bit path exists so a native render need not be converted to
// f32 whole; it has to agree with the path it replaces to within the
// quantisation it introduces.
let (w, h) = (64usize, 64usize);
let mut rgba = vec![0u8; w * h * 4];
let mut rgb = vec![0.0f32; w * h * 3];
for y in 0..h {
for x in 0..w {
let v = [
(x * 4 % 256) as u8,
(y * 4 % 256) as u8,
((x + y) % 256) as u8,
];
for c in 0..3 {
rgba[(y * w + x) * 4 + c] = v[c];
rgb[(y * w + x) * 3 + c] = v[c] as f32 / 255.0;
}
rgba[(y * w + x) * 4 + 3] = 255;
}
}
let lm = shifted_scaled(0.35, 32.0, 32.0, 0.2);
let a = warp_pixels(Pixels::RgbF32(&rgb), w, h, &lm).unwrap();
let b = warp_pixels(Pixels::Rgba8(&rgba), w, h, &lm).unwrap();
assert_eq!(a.source_px(), b.source_px());
for (x, y) in a.pixels().iter().zip(b.pixels()) {
assert!((x - y).abs() < 1e-6, "{x} vs {y}");
}
}
#[test]
fn sharpness_survives_the_contrast_being_halved() {
let edge = 200;
+48 -11
View File
@@ -136,11 +136,24 @@ pub const RIVAL_FLOOR: f32 = 0.5;
///
/// A face in no group, or one with no evidence for anybody, scores 0.
///
/// `gallery` is one flag per face — which faces may be evidence at all
/// ([`crate::embedding::MIN_GALLERY_QUALITY`]). Its length is the face count.
/// A pair is evidence *about* either face but only *from* a gallery one: a
/// probe learns from the references it matched, and a reference learns nothing
/// from a probe that happened to match it, however well. Without that, the one
/// short vector in a group would be the strongest match every face in it had.
///
/// `pairs` must be the *evidence* list — scanned at [`RIVAL_FLOOR`], not at the
/// merge threshold. Passing the merge list still works but silently removes
/// every rival weaker than a merge, which is most of them, and every uniqueness
/// collapses to 1.
pub fn identity_shares(faces: usize, clusters: &[Cluster], pairs: &[Pair], top: usize) -> Vec<f32> {
pub fn identity_shares(
gallery: &[bool],
clusters: &[Cluster],
pairs: &[Pair],
top: usize,
) -> Vec<f32> {
let faces = gallery.len();
// An identity is a *person*, not a group. One person routinely holds
// several anchored groups — the same reason they hold several unnamed ones
// — and keying this by group had Catherine competing with Catherine, which
@@ -174,15 +187,16 @@ pub fn identity_shares(faces: usize, clusters: &[Cluster], pairs: &[Pair], top:
}
// A pair is evidence in both directions: j's identity hears about i,
// and i's identity hears about j. The pair list holds each unordered
// pair once, so both have to be recorded here.
// pair once, so both have to be recorded here — each only where the
// face doing the telling is in the gallery.
let (gi, gj) = (group_of[p.i], group_of[p.j]);
if gj != usize::MAX {
if gj != usize::MAX && gallery[p.j] {
evidence[p.i]
.entry(key_of[gj])
.or_default()
.push(p.probability);
}
if gi != usize::MAX {
if gi != usize::MAX && gallery[p.i] {
evidence[p.j]
.entry(key_of[gi])
.or_default()
@@ -255,6 +269,11 @@ mod tests {
Pair { i, j, probability }
}
/// `n` faces, every one of them fit to be compared against.
fn all(n: usize) -> Vec<bool> {
vec![true; n]
}
/// The failure the module exists to fix: face 0 matches its own group's
/// three members strongly, and the group has forty more it is unrelated to.
/// The old within-group mean reported ~0.07 for this.
@@ -264,7 +283,7 @@ mod tests {
let clusters = vec![cluster(&members)];
let pairs = vec![pair(0, 1, 0.99), pair(0, 2, 0.97), pair(0, 3, 0.95)];
let shares = identity_shares(44, &clusters, &pairs, TOP_MATCHES);
let shares = identity_shares(&all(44), &clusters, &pairs, TOP_MATCHES);
assert!(
(shares[0] - 0.97).abs() < 1e-6,
"the mean of its three real matches, undiluted: {}",
@@ -284,7 +303,7 @@ mod tests {
pair(0, 4, 0.90),
];
let shares = identity_shares(5, &clusters, &pairs, TOP_MATCHES);
let shares = identity_shares(&all(5), &clusters, &pairs, TOP_MATCHES);
// Coherent at 0.90, and only half of the evidence is its own.
assert!(
(shares[0] - 0.45).abs() < 1e-6,
@@ -298,9 +317,9 @@ mod tests {
#[test]
fn a_rival_too_weak_to_merge_still_lowers_the_confidence() {
let clusters = vec![named(&[0, 1], 1), named(&[2, 3], 2)];
let sure = identity_shares(4, &clusters, &[pair(0, 1, 0.95)], TOP_MATCHES);
let sure = identity_shares(&all(4), &clusters, &[pair(0, 1, 0.95)], TOP_MATCHES);
let contested = identity_shares(
4,
&all(4),
&clusters,
&[pair(0, 1, 0.95), pair(0, 2, 0.60)],
TOP_MATCHES,
@@ -321,7 +340,7 @@ mod tests {
fn an_unnamed_group_is_not_treated_as_competition() {
let clusters = vec![cluster(&[0, 1]), cluster(&[2, 3])];
let shares = identity_shares(
4,
&all(4),
&clusters,
&[pair(0, 1, 0.95), pair(0, 2, 0.90)],
TOP_MATCHES,
@@ -343,7 +362,7 @@ mod tests {
let mut pairs: Vec<Pair> = (1..11).map(|j| pair(0, j, 0.90)).collect();
pairs.extend((11..62).map(|j| pair(0, j, 0.55)));
let shares = identity_shares(62, &clusters, &pairs, TOP_MATCHES);
let shares = identity_shares(&all(62), &clusters, &pairs, TOP_MATCHES);
// Ten at 0.90 against ten at 0.55 — not fifty-one at 0.55.
assert!(
(shares[0] - 0.90 * (9.0 / 14.5)).abs() < 1e-5,
@@ -352,11 +371,29 @@ mod tests {
);
}
/// A probe learns from the references it matched; a reference learns
/// nothing from a probe. The pair is the same pair — what differs is who
/// is doing the telling.
#[test]
fn a_face_outside_the_gallery_is_nobody_s_evidence() {
let clusters = vec![named(&[0, 1, 2], 1)];
let gallery = vec![true, true, false];
let pairs = vec![pair(0, 1, 0.80), pair(0, 2, 0.99), pair(1, 2, 0.99)];
let shares = identity_shares(&gallery, &clusters, &pairs, TOP_MATCHES);
// Faces 0 and 1 hear only from each other: the 0.99 the probe offered
// them is not counted.
assert!((shares[0] - 0.80).abs() < 1e-6, "{}", shares[0]);
assert!((shares[1] - 0.80).abs() < 1e-6, "{}", shares[1]);
// The probe hears from both references.
assert!((shares[2] - 0.99).abs() < 1e-6, "{}", shares[2]);
}
/// A face nothing has any evidence about claims nothing.
#[test]
fn a_face_with_no_evidence_reports_no_confidence() {
let clusters = vec![cluster(&[0, 1])];
let shares = identity_shares(2, &clusters, &[], TOP_MATCHES);
let shares = identity_shares(&all(2), &clusters, &[], TOP_MATCHES);
assert_eq!(shares, vec![0.0, 0.0]);
}
}
+18 -8
View File
@@ -21,11 +21,20 @@
//! fails. The exceptions (mirrors, photographs of photographs, collages) are
//! rare enough to be noise at this scale.
//!
//! **Positives have to be earned.** In order of trustworthiness: pairs the user
//! has confirmed onto one person; then burst siblings, since FR-CULL-5 already
//! groups bursts and two faces in adjacent frames are near-certainly the same
//! person. Nothing else — bootstrapping positives from high cosine is circular,
//! fitting the calibration to the belief it was supposed to test.
//! **Positives have to be earned.** A positive is a pair of faces the user has
//! confirmed onto one person (FR-CULL-10), and there is no second source: the
//! only labelling this subsystem has is the labelling somebody did by hand.
//! Bootstrapping positives from a high cosine is circular — it fits the
//! calibration to the belief it was supposed to test — and that is the whole
//! of the alternative.
//!
//! docs/faces.md §8.1 names one more that would cost no labelling at all: two
//! faces in adjacent frames of one burst are near-certainly the same person,
//! and FR-CULL-5's grouping is sitting there. Nothing draws on it. This crate
//! cannot see a catalog, let alone the bursts in one — it is handed cosines by
//! whoever assembled the pair — and no caller does that assembly on its behalf
//! yet. Until one does, and until the purity of a burst pair is *measured*
//! rather than assumed, the positives are the confirmations and nothing else.
//!
//! Which is why a fresh library has **no valid calibration** — no fit of its
//! own — and says so. It is not left without a curve: it uses the reference
@@ -44,9 +53,10 @@ const BINS: usize = 200;
/// Far stricter than the reference implementation's floor of two positives and
/// one negative. That floor is reasonable there: its pairs come from a curated
/// gallery of labelled reference portraits, where a positive pair is
/// trustworthy by construction. Here the positives are bootstrapped from bursts
/// and a handful of early confirmations, and the whole risk is fitting
/// confidently to too few of them.
/// trustworthy by construction. Here every positive is a pair somebody
/// confirmed while working through a young library's suggestions — a handful,
/// arriving slowly — and the whole risk is fitting confidently to too few of
/// them.
pub const MIN_POSITIVE_PAIRS: u64 = 200;
pub const MIN_NEGATIVE_PAIRS: u64 = 2_000;
+263
View File
@@ -0,0 +1,263 @@
//! TRACES: FR-CULL-8a
//! The two small classifiers behind a face's eye state (docs/faces.md §17).
//!
//! **OCEC** — *open closed eyes classification*, Hyodo 2025 — reads one
//! 40×24 eye and answers P(open). **SGC** — *sunglasses classification*,
//! Hyodo 2026 — reads a 48×48 head and answers P(sunglasses); it is shown
//! two framings of each face and the higher answer stands, for the reason
//! [`crate::align::SUNGLASSES_WINDOWS`] gives. Both are
//! depthwise-separable CNNs of a few hundred kilobytes, both MIT with their
//! weights, and both were exported with BatchNorm already folded, which is
//! about the friendliest graph tract can be handed.
//!
//! Neither takes a plain buffer. [`EyeClassifier::classify`] takes an
//! [`EyePatch`] and [`SunglassesClassifier::classify`] a [`HeadViews`], each
//! constructible only by the crop in [`crate::align`] that puts the right
//! pixels in it — the same defence [`crate::embed::Embedder`] makes with
//! [`crate::align::Aligned112`], for the same reason: a classifier handed the
//! wrong region returns a confident probability of nothing. Where the eye
//! box comes from is [`crate::landmarks`]; [`EyeModels::read`] is the whole
//! chain.
//!
//! # The graphs must have a fixed batch
//!
//! Both ship with a dynamic batch dimension, which tract will not analyse.
//! `tools/fix-face-model-shapes.sh` pins it to 1, exactly as it does for the
//! embedder; the shipped files are the pinned ones.
//!
//! # Pre-processing
//!
//! Read off the reference demos rather than assumed: RGB, `x / 255`, NCHW,
//! the crop resized to the input with bilinear interpolation and **without**
//! preserving its aspect. [`crate::align`]'s crops arrive already at the
//! input size in `0..=1`, so there is nothing left to do but lay them out.
use ndarray::Array4;
use crate::align::{
eye_box, eye_patch, head_views, EyePatch, HeadViews, EYE_PATCH_HEIGHT, EYE_PATCH_WIDTH,
SUNGLASSES_EDGE,
};
use crate::eyes::{Eye, EyeReading};
use crate::landmarks::{Landmarker, Landmarks};
use crate::{FaceError, Pixels};
use dr_inference_engine::{Form, Model, Role};
/// A loaded OCEC graph.
pub struct EyeClassifier {
session: Model,
}
/// A loaded SGC graph.
pub struct SunglassesClassifier {
session: Model,
}
/// Open a single-input, single-output classifier and check it is the shape
/// the crop feeding it will be.
///
/// The check is against the *input*, because that is where these two graphs
/// differ from each other and from everything else in this crate: an SGC file
/// given to the eye classifier would otherwise be resized into by an eye
/// patch, and answer. `expected` names the model in the error.
fn open_classifier(
bytes: &[u8],
expected: &'static str,
(h, w): (usize, usize),
) -> Result<Model, FaceError> {
let model = dr_inference_engine::open(Role::EyeClassifier, Form::F32, bytes)?;
let acquired = model.acquire()?;
let session = acquired.lock();
let input = session.inputs().first().ok_or(FaceError::WrongModel {
expected,
detail: "model has no inputs".into(),
})?;
let shape: Option<Vec<i64>> = input.dtype().tensor_shape().map(|s| s.to_vec());
let want = [1, 3, h as i64, w as i64];
if shape.as_deref() != Some(&want[..]) {
return Err(FaceError::WrongModel {
expected,
detail: format!(
"input '{}' is {:?}, expected {:?} (batch pinned to 1)",
input.name(),
shape,
want
),
});
}
if session.outputs().len() != 1 {
return Err(FaceError::WrongModel {
expected,
detail: format!("{} outputs, expected one", session.outputs().len()),
});
}
drop(session);
drop(acquired);
Ok(model)
}
/// Lay a `h × w` RGB crop out as the `[1, 3, h, w]` tensor both graphs take.
fn to_nchw(pixels: &[f32], h: usize, w: usize) -> Array4<f32> {
let mut input = Array4::<f32>::zeros((1, 3, h, w));
for y in 0..h {
for x in 0..w {
for c in 0..3 {
input[[0, c, y, x]] = pixels[(y * w + x) * 3 + c];
}
}
}
input
}
/// Run a one-number classifier and read its sigmoid back, clamped.
fn run_scalar(model: &Model, input: Array4<f32>, expected: &'static str) -> Result<f32, FaceError> {
let acquired = model.acquire()?;
let mut session = acquired.lock();
let outputs = session
.run(ort::inputs![
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
])
.map_err(FaceError::Inference)?;
let (_, data) = outputs[0]
.try_extract_tensor::<f32>()
.map_err(FaceError::Inference)?;
let Some(&p) = data.first() else {
return Err(FaceError::WrongModel {
expected,
detail: "empty output".into(),
});
};
// The graph ends in a sigmoid, so this is a clamp against rounding and
// nothing more — the reference demo does the same.
Ok(p.clamp(0.0, 1.0))
}
impl EyeClassifier {
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
Self::from_bytes(&bytes)
}
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
Ok(Self {
session: open_classifier(bytes, "OCEC", (EYE_PATCH_HEIGHT, EYE_PATCH_WIDTH))?,
})
}
/// P(open) for one eye.
pub fn classify(&mut self, eye: &EyePatch) -> Result<f32, FaceError> {
let input = to_nchw(eye.pixels(), EYE_PATCH_HEIGHT, EYE_PATCH_WIDTH);
run_scalar(&self.session, input, "OCEC")
}
}
impl SunglassesClassifier {
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
Self::from_bytes(&bytes)
}
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
Ok(Self {
session: open_classifier(bytes, "SGC", (SUNGLASSES_EDGE, SUNGLASSES_EDGE))?,
})
}
/// P(sunglasses) for one head: the highest answer over its framings.
pub fn classify(&mut self, head: &HeadViews) -> Result<f32, FaceError> {
let mut best = 0.0_f32;
for view in head.views() {
let input = to_nchw(view, SUNGLASSES_EDGE, SUNGLASSES_EDGE);
best = best.max(run_scalar(&self.session, input, "SGC")?);
}
Ok(best)
}
}
/// The three models behind a reading, which is how every caller holds them.
///
/// One struct rather than three optional parameters, because a partial
/// reading is not a reading: an eye state with no sunglasses number behind
/// it is exactly the beach-photograph failure [`crate::eyes`] describes, and
/// an eye box without the landmarks is the loose one this module replaced.
/// The models load together or not at all.
pub struct EyeModels {
pub landmarks: Landmarker,
pub eyes: EyeClassifier,
pub sunglasses: SunglassesClassifier,
}
impl EyeModels {
pub fn from_paths(
landmarks: impl AsRef<std::path::Path>,
eyes: impl AsRef<std::path::Path>,
sunglasses: impl AsRef<std::path::Path>,
) -> Result<Self, FaceError> {
Ok(Self {
landmarks: Landmarker::from_path(landmarks)?,
eyes: EyeClassifier::from_path(eyes)?,
sunglasses: SunglassesClassifier::from_path(sunglasses)?,
})
}
/// Read one face's eyes, and hand back the dense landmarks it read them
/// from.
///
/// `bbox` is the detector's `(x0, y0, x1, y1)` and `landmarks5` its five
/// points, both in source pixels; the buffer is the one the aligned
/// crop was taken from, so an eye is read from the same pixels the
/// embedder saw the face in. `None` where nothing could be cut — a
/// degenerate box or landmarks — which the caller stores as "not read".
///
/// The landmarks come back because they cost a model run the caller will
/// not want to pay twice: stored beside the reading, a later pass over
/// faces — head pose, expression — has them without the original.
pub fn read(
&mut self,
px: Pixels<'_>,
width: usize,
height: usize,
bbox: (f32, f32, f32, f32),
landmarks5: &[(f32, f32); 5],
) -> Result<Option<(EyeReading, Landmarks)>, FaceError> {
let Some(lm) = self.landmarks.landmarks(px, width, height, bbox)? else {
return Ok(None);
};
let Some(head) = head_views(px, width, height, landmarks5) else {
return Ok(None);
};
let mut eye = |contour: &[(f32, f32)]| -> Result<Eye, FaceError> {
// A hidden eye's contour can collapse to no width. Its numbers
// are then zero — no pixels, no sharpness — which is what the
// rule in `crate::eyes` reads as "not readable".
let Some(b) = eye_box(contour) else {
return Ok(Eye {
open: 0.0,
px: 0.0,
sharpness: 0.0,
});
};
let Some(patch) = eye_patch(px, width, height, b) else {
return Ok(Eye {
open: 0.0,
px: 0.0,
sharpness: 0.0,
});
};
Ok(Eye {
open: self.eyes.classify(&patch)?,
px: patch.source_px(),
sharpness: patch.sharpness(),
})
};
let right = eye(&lm.right_eye())?;
let left = eye(&lm.left_eye())?;
let reading = EyeReading {
right,
left,
sunglasses: self.sunglasses.classify(&head)?,
};
Ok(Some((reading, lm)))
}
}
+307 -3
View File
@@ -19,6 +19,23 @@
//! and clustering never moves it. Two groups holding confirmations of
//! *different* people cannot merge, whatever their similarity says.
//!
//! # The gallery, and the faces that are only ever compared against it
//!
//! A third defence, and the cheapest of all: **a short embedding is never a
//! reference.** The length of the raw vector is the model's own reading of
//! how recognisable the crop was ([`crate::embedding::MIN_GALLERY_QUALITY`]),
//! and a short one sits near the centre of the sphere, matching a little of
//! everybody. One of those in a group is a bridge to the next group over.
//!
//! So the population is split. Faces at or above the floor are the
//! **gallery**, and they cluster exactly as described below. Faces under it
//! are **probes**: each is measured against the finished groups and joins the
//! one it fits, by the same average-link rule and under the same constraints
//! — but it is measured against the gallery members only, never against
//! another probe, and once placed it is never part of what the next face is
//! measured against. A blurred photograph of a known person is still named;
//! it just cannot vouch for anyone else.
//!
//! # Average link, not single link
//!
//! Single-link chains: one bad edge welds two identities together, and it is
@@ -117,6 +134,11 @@ pub struct Candidate {
pub embedding: Vec<f32>,
/// Source pixels across the aligned crop, for the calibration's size term.
pub crop_px: f32,
/// Length of the raw embedding, where it was recorded
/// ([`crate::embedding::MIN_GALLERY_QUALITY`]). `None` for a face indexed
/// before it was kept, which is admitted to the gallery — see
/// [`Candidate::in_gallery`].
pub quality: Option<f32>,
/// The person this face is *confirmed* to be, if any.
///
/// Suggestions are deliberately not passed here. They are this function's
@@ -125,6 +147,13 @@ pub struct Candidate {
pub confirmed_person: Option<u64>,
}
impl Candidate {
/// Whether this face may be compared *against*, as well as compared.
pub fn in_gallery(&self) -> bool {
crate::embedding::in_gallery(self.quality)
}
}
/// One group of faces the clusterer believes are one person.
#[derive(Debug, Clone, PartialEq)]
pub struct Cluster {
@@ -202,7 +231,7 @@ pub fn cluster_scored(faces: &[Candidate], cal: &Calibration, min_probability: f
let clusters = build(faces, cal, min_probability, &merges);
let confidence = crate::assign::identity_shares(
faces.len(),
&columns.gallery,
&clusters,
&evidence,
crate::assign::TOP_MATCHES,
@@ -225,6 +254,7 @@ struct Columns {
dim: usize,
crop_px: Vec<f32>,
images: Vec<u64>,
gallery: Vec<bool>,
}
impl Columns {
@@ -245,6 +275,7 @@ impl Columns {
dim,
crop_px: faces.iter().map(|f| f.crop_px).collect(),
images: faces.iter().map(|f| f.image).collect(),
gallery: faces.iter().map(Candidate::in_gallery).collect(),
}
}
@@ -254,22 +285,171 @@ impl Columns {
dim: self.dim,
crop_px: &self.crop_px,
images: &self.images,
gallery: &self.gallery,
}
}
}
/// Agglomerate the gallery over its pairs, then place the probes.
///
/// `pairs` is what [`neighbours::above_threshold`] returned: every pair has a
/// gallery side, but a pair with a probe on the other side is not a merge —
/// it is the evidence [`place_probes`] works from. Only the gallery-to-gallery
/// pairs reach the engine, so a probe enters it as a singleton with no edges
/// and comes out exactly as it went in.
fn build(
faces: &[Candidate],
cal: &Calibration,
min_probability: f32,
pairs: &[neighbours::Pair],
) -> Vec<Cluster> {
let gallery: Vec<bool> = faces.iter().map(Candidate::in_gallery).collect();
let (merges, probe_pairs): (Vec<_>, Vec<_>) = pairs
.iter()
.copied()
.partition(|p| gallery[p.i] && gallery[p.j]);
let mut engine = Engine::new(faces, cal, min_probability);
let parts = components(faces.len(), pairs);
let parts = components(faces.len(), &merges);
for (component, edges) in parts.members.iter().zip(&parts.edges) {
engine.agglomerate(component, edges);
}
engine.finish()
let dot = engine.dot;
let clusters = engine.finish();
if probe_pairs.is_empty() {
return clusters;
}
place_probes(
faces,
cal,
min_probability,
dot,
&gallery,
clusters,
&probe_pairs,
)
}
/// Put each probe into the finished group it fits, or leave it alone.
///
/// The same decision the engine makes for a singleton — average link over the
/// group, at or above `min_probability`, subject to [`Engine::can_link`]'s two
/// constraints — with one difference that is the whole point: the average is
/// over the group's **gallery** members. A probe already placed is not part of
/// what the next one is measured against, so a run of short vectors cannot
/// pull each other in one after another.
///
/// Probes are placed in index order and each placement is final, which is
/// what keeps this deterministic. The group a probe joins gains its
/// photograph, so a second face from the same frame cannot follow it — the
/// co-occurrence rule, applied exactly as the engine applies it.
fn place_probes(
faces: &[Candidate],
cal: &Calibration,
min_probability: f32,
dot: neighbours::DotFn,
gallery: &[bool],
mut clusters: Vec<Cluster>,
probe_pairs: &[neighbours::Pair],
) -> Vec<Cluster> {
// Where each face sits, and what each group's photographs and gallery
// members are. The probe's own singleton is here too, and is dropped once
// it has moved.
let mut group_of = vec![usize::MAX; faces.len()];
for (g, c) in clusters.iter().enumerate() {
for &m in &c.members {
group_of[m] = g;
}
}
let mut images: Vec<HashSet<u64>> = clusters
.iter()
.map(|c| c.members.iter().map(|&m| faces[m].image).collect())
.collect();
let references: Vec<Vec<usize>> = clusters
.iter()
.map(|c| c.members.iter().copied().filter(|&m| gallery[m]).collect())
.collect();
// Which groups each probe has any above-threshold pair into. Only those
// can average above the threshold — the argument the module note makes
// for the engine holds here unchanged.
let mut candidates: Vec<Vec<usize>> = vec![Vec::new(); faces.len()];
for p in probe_pairs {
let (probe, reference) = if gallery[p.i] { (p.j, p.i) } else { (p.i, p.j) };
candidates[probe].push(group_of[reference]);
}
let mut moved: Vec<usize> = Vec::new();
for probe in 0..faces.len() {
if gallery[probe] || candidates[probe].is_empty() {
continue;
}
let mut groups = std::mem::take(&mut candidates[probe]);
groups.sort_unstable();
groups.dedup();
let face = &faces[probe];
let mut best: Option<(f32, usize)> = None;
for g in groups {
let target = &clusters[g];
if let (Some(mine), Some(theirs)) = (face.confirmed_person, target.person) {
if mine != theirs {
continue;
}
}
if images[g].contains(&face.image) {
continue;
}
let (mut sum, mut count) = (0.0_f64, 0.0_f64);
for &r in &references[g] {
let cos = dot(&face.embedding, &faces[r].embedding);
let min_crop = face.crop_px.min(faces[r].crop_px);
sum += cal.probability(cos, min_crop, 0.0) as f64;
count += 1.0;
}
if count == 0.0 {
continue;
}
let p = (sum / count) as f32;
// Strictly better wins; on a tie the lowest group index, which is
// the engine's own tiebreak.
if p >= min_probability && best.is_none_or(|(bp, _)| p > bp) {
best = Some((p, g));
}
}
let Some((_, g)) = best else { continue };
let own = group_of[probe];
clusters[g].members.push(probe);
clusters[g].members.sort_unstable();
clusters[g].person = clusters[g].person.or(face.confirmed_person);
images[g].insert(face.image);
group_of[probe] = g;
moved.push(own);
}
if moved.is_empty() {
return clusters;
}
// The singletons the probes left behind, then the order `Engine::finish`
// promises: largest first, lowest member first among equals.
let mut vacated = vec![false; clusters.len()];
for g in moved {
vacated[g] = true;
}
let mut out: Vec<Cluster> = clusters
.into_iter()
.zip(vacated)
.filter(|(_, gone)| !gone)
.map(|(c, _)| c)
.collect();
out.sort_by(|x, y| {
y.members
.len()
.cmp(&x.members.len())
.then(x.members[0].cmp(&y.members[0]))
});
out
}
/// Split one person's faces into the groups a raised threshold separates them
@@ -726,10 +906,19 @@ mod tests {
image,
embedding: at_cosine(identity, cosine),
crop_px: 150.0,
quality: None,
confirmed_person: None,
}
}
/// A face too short to be a reference: compared, never compared against.
fn probe(face: u64, image: u64, identity: usize, cosine: f32) -> Candidate {
Candidate {
quality: Some(crate::embedding::MIN_GALLERY_QUALITY - 5.0),
..candidate(face, image, identity, cosine)
}
}
/// A calibration steep enough that the test's cosines are unambiguous:
/// 0.6 is near-certain, 0.1 is near-impossible.
fn cal() -> Calibration {
@@ -1099,6 +1288,7 @@ mod tests {
image,
embedding: at_cosine(p, cosine),
crop_px: 60.0 + ((out.len() % 11) as f32) * 25.0,
quality: None,
confirmed_person: None,
});
image += 1;
@@ -1182,6 +1372,7 @@ mod tests {
image: 5_000,
embedding: at_cosine(200, 1.0),
crop_px: 150.0,
quality: None,
confirmed_person: None,
});
let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
@@ -1191,4 +1382,117 @@ mod tests {
"the outlier was absorbed"
);
}
// ── the gallery ───────────────────────────────────────────────────────
/// A short vector is still somebody: it joins the group it matches.
#[test]
fn a_probe_joins_the_group_it_matches() {
let faces = vec![
candidate(1, 10, 0, 1.0),
candidate(2, 11, 0, 0.95),
probe(3, 12, 0, 0.92),
];
let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
assert_eq!(out.len(), 1);
assert_eq!(out[0].members, vec![0, 1, 2]);
}
/// Two short vectors that resemble each other are noise agreeing with
/// noise, and there is nothing in the gallery for either to be measured
/// against.
#[test]
fn two_probes_are_never_grouped_with_each_other() {
let faces = vec![probe(1, 10, 0, 1.0), probe(2, 11, 0, 0.98)];
let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
assert_eq!(out.len(), 2, "two probes were grouped: {out:?}");
}
/// The point of measuring against the gallery only: a probe that has been
/// placed is not a stepping stone for the next one.
#[test]
fn a_placed_probe_is_not_what_the_next_probe_is_measured_against() {
let mut first = probe(2, 11, 0, 0.6);
// 0.6 along identity 0 and 0.8 along its perpendicular: near enough to
// the reference to join it, and much nearer to the face below.
first.embedding = at_cosine(0, 0.6);
let mut second = probe(3, 12, 0, 0.0);
second.embedding = at_cosine(0, 0.0);
let faces = vec![candidate(1, 10, 0, 1.0), first, second];
let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
let group = out.iter().find(|c| c.members.contains(&0)).unwrap();
assert_eq!(
group.members,
vec![0, 1],
"the first probe should have joined"
);
assert!(
out.iter().any(|c| c.members == vec![2]),
"the second probe reached the group through the first: {out:?}"
);
}
/// A confirmation on a probe is still the user's word: the group it joins
/// becomes that person, and a group already someone else's is closed to it.
#[test]
fn a_probe_carries_its_confirmation_and_respects_others() {
let mut anchored = probe(3, 12, 0, 0.92);
anchored.confirmed_person = Some(7);
let faces = vec![
candidate(1, 10, 0, 1.0),
candidate(2, 11, 0, 0.95),
anchored,
];
let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
assert_eq!(out.len(), 1);
assert_eq!(out[0].person, Some(7));
let mut theirs = candidate(1, 10, 0, 1.0);
theirs.confirmed_person = Some(8);
let faces = vec![theirs, candidate(2, 11, 0, 0.95), {
let mut a = probe(3, 12, 0, 0.92);
a.confirmed_person = Some(7);
a
}];
let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
assert!(
out.iter()
.any(|c| c.members == vec![2] && c.person == Some(7)),
"a probe confirmed as one person joined another's group: {out:?}"
);
}
/// The co-occurrence rule follows a probe in: once it has joined, its
/// photograph is the group's.
#[test]
fn a_probe_cannot_join_a_group_holding_a_face_from_its_own_photograph() {
let faces = vec![
candidate(1, 10, 0, 1.0),
candidate(2, 11, 0, 0.95),
probe(3, 10, 0, 0.92),
];
let out = cluster(&faces, &cal(), DEFAULT_MERGE_PROBABILITY);
assert!(out.iter().any(|c| c.members == vec![2]), "{out:?}");
}
/// A probe's placement is scored like anyone else's, from the references
/// it matched — and the references' own scores do not hear from it.
#[test]
fn a_probe_is_scored_but_is_not_evidence() {
let gallery_only = vec![candidate(1, 10, 0, 1.0), candidate(2, 11, 0, 0.95)];
let without = cluster_scored(&gallery_only, &cal(), DEFAULT_MERGE_PROBABILITY);
let mut with_probe = gallery_only.clone();
with_probe.push(probe(3, 12, 0, 0.99));
let with = cluster_scored(&with_probe, &cal(), DEFAULT_MERGE_PROBABILITY);
assert_eq!(with.clusters[0].members, vec![0, 1, 2]);
assert!(with.confidence[2] > 0.9, "{}", with.confidence[2]);
assert_eq!(
&with.confidence[..2],
&without.confidence[..],
"a probe changed what the references were sure of"
);
}
}
+36 -14
View File
@@ -14,7 +14,8 @@
use ndarray::Array4;
use crate::{install_backend, FaceError};
use crate::FaceError;
use dr_inference_engine::{Form, Model, Role};
/// The graph's input edge, in pixels. See the module note: not configurable.
pub const INPUT_EDGE: usize = 640;
@@ -135,7 +136,10 @@ impl Detection {
/// A loaded SCRFD graph.
pub struct Detector {
session: ort::session::Session,
session: Model,
/// f32 or int8 — the int8 form finds a different set of faces and is a
/// different detector in `model_id` (docs/inference.md §7).
form: Form,
/// Feature-map count: 3 for strides {8,16,32}, 4 for {8,16,32,64}.
///
/// Discovered from the output count rather than assumed, because both
@@ -145,18 +149,29 @@ pub struct Detector {
}
impl Detector {
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
Self::from_bytes(&bytes)
/// Which form this detector was loaded from.
pub fn form(&self) -> Form {
self.form
}
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
install_backend();
/// Load the canonical f32 file at `path`, or the form the device's
/// backend wants instead — the `.int8.onnx` beside it on a Hexagon —
/// which [`Detector::form`] then reports.
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let (path, form) = dr_inference_engine::resolve_model(Role::Detector, path.as_ref());
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
Self::from_bytes_in(&bytes, form)
}
let session = ort::session::Session::builder()
.map_err(FaceError::Inference)?
.commit_from_memory(bytes)
.map_err(FaceError::Inference)?;
/// An f32 graph from memory.
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
Self::from_bytes_in(bytes, Form::F32)
}
fn from_bytes_in(bytes: &[u8], form: Form) -> Result<Self, FaceError> {
let model = dr_inference_engine::open(Role::Detector, form, bytes)?;
let acquired = model.acquire()?;
let session = acquired.lock();
let n_out = session.outputs().len();
if n_out % 3 != 0 || !(9..=12).contains(&n_out) {
@@ -191,7 +206,13 @@ impl Detector {
}
}
Ok(Self { session, fmc })
drop(session);
drop(acquired);
Ok(Self {
session: model,
form,
fmc,
})
}
/// Stride levels this graph emits.
@@ -223,8 +244,9 @@ impl Detector {
let lb = Letterbox::fit(width as f32, height as f32);
let input = lb.sample(rgb, width, height);
let outputs = self
.session
let acquired = self.session.acquire()?;
let mut session = acquired.lock();
let outputs = session
.run(ort::inputs![
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
])
+60 -16
View File
@@ -14,11 +14,47 @@ use ndarray::Array4;
use crate::align::{Aligned112, ALIGNED_EDGE};
use crate::embedding::{normalise, Embedding, ModelId, EMBEDDING_DIM};
use crate::{install_backend, FaceError};
use crate::FaceError;
use dr_inference_engine::{Form, Model, Role};
/// What one pass of the embedder produces: the direction, and the length.
///
/// Two fields rather than a `quality` on [`Embedding`], because every other
/// holder of an `Embedding` relies on it being unit length and compares by
/// dot product; the length is a separate fact about the same face, and it is
/// stored separately too.
#[derive(Debug, Clone, PartialEq)]
pub struct Embedded {
pub embedding: Embedding,
/// L2 norm of the raw model output.
///
/// The model's own opinion of how recognisable the crop was — see
/// [`crate::embedding::MIN_GALLERY_QUALITY`] for what it means and where
/// it is used.
pub quality: f32,
}
impl Embedded {
/// Storage form: the **raw** vector, `512 × f16`.
///
/// Not the unit vector. The length is the quality, and a store that held
/// only the direction would have thrown it away at the one moment it could
/// be known — which is what this crate used to do. Readers re-normalise
/// ([`Embedding::from_f16_bytes`]), so every comparison is still a dot
/// product, and [`crate::embedding::read_f16_bytes`] gives the length back
/// to a reader that wants it.
///
/// f16 costs nothing extra at this scale: its precision is relative, so a
/// component of a vector of length 20 is kept to the same three figures as
/// the same component scaled to length 1.
pub fn to_f16_bytes(&self) -> Vec<u8> {
self.embedding.to_f16_bytes_scaled(self.quality)
}
}
/// A loaded ArcFace graph.
pub struct Embedder {
session: ort::session::Session,
session: Model,
model: ModelId,
}
@@ -29,12 +65,11 @@ impl Embedder {
}
pub fn from_bytes(bytes: &[u8], model: ModelId) -> Result<Self, FaceError> {
install_backend();
let session = ort::session::Session::builder()
.map_err(FaceError::Inference)?
.commit_from_memory(bytes)
.map_err(FaceError::Inference)?;
// Always the f32 form: an embedding must compare across devices
// (docs/inference.md §7), and the engine pins this role to it.
let loaded = dr_inference_engine::open(Role::Embedder, Form::F32, bytes)?;
let acquired = loaded.acquire()?;
let session = acquired.lock();
// One output, `[1, 512]`. Checked because an ArcFace variant with a
// different embedding width would otherwise be read as a truncated
@@ -55,7 +90,12 @@ impl Embedder {
});
}
Ok(Self { session, model })
drop(session);
drop(acquired);
Ok(Self {
session: loaded,
model,
})
}
pub fn model(&self) -> &ModelId {
@@ -63,7 +103,7 @@ impl Embedder {
}
/// Embed one aligned face.
pub fn embed(&mut self, face: &Aligned112) -> Result<Embedding, FaceError> {
pub fn embed(&mut self, face: &Aligned112) -> Result<Embedded, FaceError> {
// `(x·255 − 127.5) / 128` — see the `/128` note in `detect::Letterbox`.
let px = face.pixels();
let mut input = Array4::<f32>::zeros((1, 3, ALIGNED_EDGE, ALIGNED_EDGE));
@@ -76,8 +116,9 @@ impl Embedder {
}
}
let outputs = self
.session
let acquired = self.session.acquire()?;
let mut session = acquired.lock();
let outputs = session
.run(ort::inputs![
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
])
@@ -95,11 +136,14 @@ impl Embedder {
let mut v = Box::new([0.0_f32; EMBEDDING_DIM]);
v.copy_from_slice(&data[..EMBEDDING_DIM]);
normalise(&mut v);
let quality = normalise(&mut v);
Ok(Embedding {
model: self.model.clone(),
v,
Ok(Embedded {
embedding: Embedding {
model: self.model.clone(),
v,
},
quality,
})
}
}
+119 -18
View File
@@ -12,6 +12,43 @@
/// Embedding dimensionality. Fixed by the model family, not a parameter.
pub const EMBEDDING_DIM: usize = 512;
/// The shortest raw embedding a face may be *compared against*.
///
/// # What the length of the vector says
///
/// ArcFace is trained on the direction of its output and nothing else, and
/// the length it leaves behind turns out to be a free quality signal: the
/// magnitude grows with how recognisable the crop was to the model, and a
/// blurred, occluded, badly lit or hard-profile face comes out short. MagFace
/// (Meng et al., CVPR 2021) made that the training objective; the plain
/// ArcFace heads this crate runs already show it, weaker but usable, which is
/// why it is worth keeping the number the normalisation discards.
///
/// # Why it gates the gallery and not the face
///
/// A short vector is a bad *reference*: it sits nearer the centre of the
/// sphere than a real identity does and matches a little of everyone, which
/// is exactly the face that welds two people together in a clustering pass.
/// It is not a bad *probe* — the face is still real, still somebody, and
/// comparing it against good references is the only way it will ever be named.
/// So a face below this floor is compared against the gallery and never
/// becomes part of it: see `cluster::Candidate::in_gallery`.
///
/// 14 is the operating point for `w600k_mbf`, whose norms on the reference
/// library run from about 8 on a blur to the high 20s on a clean portrait. A
/// face whose quality was never recorded — indexed before the number was kept
/// — is not gated, because a rule that cannot be checked should admit, not
/// exclude.
pub const MIN_GALLERY_QUALITY: f32 = 14.0;
/// Whether an embedding of this quality may serve as a reference.
///
/// `None` is "not measured", and is admitted: the rule is about a number that
/// was read and found short, not about a number that is missing.
pub fn in_gallery(quality: Option<f32>) -> bool {
quality.is_none_or(|q| q >= MIN_GALLERY_QUALITY)
}
/// Which model produced an embedding.
///
/// Embeddings from different models are not comparable, and this is the one
@@ -58,10 +95,19 @@ impl Embedding {
}
/// Storage form: `512 × f16`, 1 KB per face (catalog.md §10.1).
///
/// This writes the unit vector. What the catalog stores is the raw one —
/// `embed::Embedded::to_f16_bytes` — because the length is the quality
/// and a unit vector has none left to read.
pub fn to_f16_bytes(&self) -> Vec<u8> {
self.to_f16_bytes_scaled(1.0)
}
/// The unit vector scaled by `length`, as `512 × f16`.
pub(crate) fn to_f16_bytes_scaled(&self, length: f32) -> Vec<u8> {
let mut out = Vec::with_capacity(EMBEDDING_DIM * 2);
for &x in self.v.iter() {
out.extend_from_slice(&f32_to_f16_bits(x).to_le_bytes());
out.extend_from_slice(&f32_to_f16_bits(x * length).to_le_bytes());
}
out
}
@@ -71,25 +117,42 @@ impl Embedding {
/// The f16 round-trip perturbs a unit vector by ~1e-3 in cosine — three
/// orders below the separation between a match and a non-match — but the
/// drift is free to remove and invisible if left, so it is removed here
/// rather than remembered at every call site.
/// rather than remembered at every call site. The same pass is what turns
/// a stored raw vector back into the unit one every comparison expects.
pub fn from_f16_bytes(model: ModelId, bytes: &[u8]) -> Option<Self> {
if bytes.len() != EMBEDDING_DIM * 2 {
return None;
}
let mut v = Box::new([0.0_f32; EMBEDDING_DIM]);
for (i, chunk) in bytes.chunks_exact(2).enumerate() {
v[i] = f16_bits_to_f32(u16::from_le_bytes([chunk[0], chunk[1]]));
}
normalise(&mut v);
Some(Self { model, v })
read_f16_bytes(model, bytes).map(|(e, _)| e)
}
}
/// Read a stored vector back, with the length it was stored at.
///
/// The length is the quality where the blob is a raw one, and ~1 where it is
/// a unit vector from before raw vectors were stored — which is why the
/// catalog keeps the quality beside the blob rather than deriving it from
/// this: a unit vector reads as a quality of 1, not as "unmeasured".
pub fn read_f16_bytes(model: ModelId, bytes: &[u8]) -> Option<(Embedding, f32)> {
if bytes.len() != EMBEDDING_DIM * 2 {
return None;
}
let mut v = Box::new([0.0_f32; EMBEDDING_DIM]);
for (i, chunk) in bytes.chunks_exact(2).enumerate() {
v[i] = f16_bits_to_f32(u16::from_le_bytes([chunk[0], chunk[1]]));
}
let length = normalise(&mut v);
Some((Embedding { model, v }, length))
}
fn dot(a: &[f32; EMBEDDING_DIM], b: &[f32; EMBEDDING_DIM]) -> f32 {
a.iter().zip(b.iter()).map(|(x, y)| x * y).sum()
}
pub(crate) fn normalise(v: &mut [f32; EMBEDDING_DIM]) {
/// Scale `v` to unit length, and return the length it had.
///
/// The length is the one thing about the raw output that survives being
/// thrown away by everything downstream, and it is a quality signal
/// ([`MIN_GALLERY_QUALITY`]) — so it comes back out rather than being lost
/// here.
pub(crate) fn normalise(v: &mut [f32; EMBEDDING_DIM]) -> f32 {
// Clamped rather than checked: a zero-norm embedding is a broken model,
// not a runtime condition worth an error path, and dividing by 1e-6 keeps
// the NaN out of the catalog.
@@ -97,6 +160,7 @@ pub(crate) fn normalise(v: &mut [f32; EMBEDDING_DIM]) {
for x in v.iter_mut() {
*x /= norm;
}
norm
}
// ── f16 ───────────────────────────────────────────────────────────────────
@@ -113,9 +177,10 @@ fn f32_to_f16_bits(x: f32) -> u16 {
let mant = bits & 0x007f_ffff;
if exp >= 0x1f {
// Overflow, inf, or NaN. Embeddings are unit-norm so this is the
// broken-model path; infinity is the honest answer, not a clamp that
// hides it.
// Overflow, inf, or NaN. No component of an embedding exceeds its
// length, and the lengths this model produces are in the tens, so
// this is the broken-model path; infinity is the honest answer, not a
// clamp that hides it.
return sign
| 0x7c00
| if mant != 0 && exp == 0x1f + 112 {
@@ -125,9 +190,9 @@ fn f32_to_f16_bits(x: f32) -> u16 {
};
}
if exp <= 0 {
// Subnormal or underflow. A component of a unit 512-vector is ~0.04,
// nowhere near here, so this branch exists for correctness rather than
// for traffic.
// Subnormal or underflow. A component of a unit 512-vector is ~0.04
// and a stored one is that times the length, nowhere near here, so
// this branch exists for correctness rather than for traffic.
if exp < -10 {
return sign;
}
@@ -221,6 +286,42 @@ mod tests {
}
}
#[test]
fn normalising_reports_the_length_it_removed() {
let mut v = Box::new([0.0_f32; EMBEDDING_DIM]);
v[0] = 3.0;
v[1] = 4.0;
let norm = normalise(&mut v);
assert!((norm - 5.0).abs() < 1e-6, "norm {norm}");
assert!((v[0] - 0.6).abs() < 1e-6 && (v[1] - 0.8).abs() < 1e-6);
}
/// The gate admits what it cannot measure: a face from before the number
/// was kept is not a face that was found wanting.
#[test]
fn an_unmeasured_quality_is_admitted_to_the_gallery() {
assert!(in_gallery(None));
assert!(in_gallery(Some(MIN_GALLERY_QUALITY)));
assert!(in_gallery(Some(27.5)));
assert!(!in_gallery(Some(MIN_GALLERY_QUALITY - 0.01)));
assert!(!in_gallery(Some(8.0)));
}
/// The storage form carries the length, and the length comes back out —
/// without touching the direction every comparison is made on.
#[test]
fn a_raw_vector_round_trips_with_its_length() {
let e = unit(3);
let raw = e.to_f16_bytes_scaled(21.5);
let (back, length) = read_f16_bytes(e.model.clone(), &raw).unwrap();
assert!((length - 21.5).abs() < 0.05, "length {length}");
assert!(e.cosine(&back).unwrap() > 0.9999);
// A unit vector from an older store reads as length 1, not as an
// error — see `read_f16_bytes` on why that is not "unmeasured".
let (_, one) = read_f16_bytes(e.model.clone(), &e.to_f16_bytes()).unwrap();
assert!((one - 1.0).abs() < 1e-2, "length {one}");
}
#[test]
fn f16_round_trip_rejects_a_wrong_length_blob() {
assert!(Embedding::from_f16_bytes(ModelId::new("m"), &[0u8; 100]).is_none());
+263
View File
@@ -0,0 +1,263 @@
//! TRACES: FR-CULL-8a
//! What a face's eyes are doing, and how the numbers behind it are read.
//!
//! Model-free: the models in [`crate::classify`] produce the numbers, and
//! everything that interprets them — the catalog's filter, the People
//! screen's label — comes through here, so a threshold lives in exactly one
//! place.
//!
//! # Seven numbers, one answer
//!
//! An eye classifier answers "open or closed" for whatever it is shown, and
//! it is shown three things it cannot answer for. **Dark glass**: over
//! sunglasses it answers anyway, confidently, for a state that cannot be
//! seen — so the reading carries P(sunglasses) from a classifier that looks
//! at the whole head, and that takes precedence. **A smear**: a soft eye is
//! not a closed one, but shown a blur the classifier says "closed" with the
//! same confidence it says anything, and on the reference library that was
//! the commonest wrong answer of all — small faces, motion, a proxy where
//! the native render should have been. So each eye carries how many source
//! pixels it spanned and how sharp the patch was, and an eye under either
//! floor is not asked. **A cheek**: a head turned far enough hides its far
//! eye, and the landmark contour of a hidden eye collapses to a sliver; an
//! eye much narrower than its partner is not asked either.
//!
//! The two eyes are kept apart rather than averaged. A wink is one eye
//! closed, and averaging it lands at 0.5 — the one value that says the least.
//! [`EyeState::Open`] requires every eye that *could be read* to be open;
//! a face with no readable eye is [`EyeState::Unreadable`], which is not a
//! blink and not open, and a filter for either leaves it alone.
/// One eye's numbers.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Eye {
/// P(open), the classifier's sigmoid.
pub open: f32,
/// Source pixels across the eye box — [`crate::align::EyePatch::source_px`].
pub px: f32,
/// [`crate::align::EyePatch::sharpness`] of the patch the classifier saw.
pub sharpness: f32,
}
/// The numbers the models produced for one face.
///
/// Stored per face, nullable as a whole: a face indexed before the eye models
/// existed, or on a device without them, has no reading rather than a
/// reading of zeros.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct EyeReading {
/// The subject's **right** eye — image-left.
pub right: Eye,
/// The subject's **left** eye — image-right.
pub left: Eye,
/// P(the head wears sunglasses).
pub sunglasses: f32,
}
/// Above this an eye is open. The classifier's own decision point; its
/// training put the two classes either side of a sigmoid and this is where
/// the sigmoid crosses.
pub const EYES_OPEN_THRESHOLD: f32 = 0.5;
/// Above this the head wears sunglasses and the eye readings are moot.
pub const SUNGLASSES_THRESHOLD: f32 = 0.5;
/// Fewest source pixels across an eye box for the eye to be read.
///
/// The classifier was trained on eyes down to about a dozen pixels wide
/// (its reference footage averaged 15–21); below that the 40-pixel patch is
/// an interpolation of nothing, and the answer is noise that reads as
/// "closed". docs/faces.md §17.3 has the measurement behind the number.
pub const MIN_EYE_PX: f32 = 12.0;
/// Least [`Eye::sharpness`] for the eye to be read.
///
/// The same measure as the face's `min_sharpness`, over the eye patch, and
/// chosen the same way: the value under which the open-eyed faces of the
/// reference sample were being called closed. docs/faces.md §17.3.
pub const MIN_EYE_SHARPNESS: f32 = 0.02;
/// An eye narrower than this fraction of its partner is the far eye of a
/// turned head, out of view behind the nose, and is not read.
///
/// A landmark model's contour for a hidden eye collapses towards the nose.
/// Measured on twenty native renders of the reference library
/// (docs/faces.md §17.4): profiles put the far eye at 0.02–0.43 of the near
/// one, two three-quarter faces whose far eye read closed sat at 0.54, and
/// every face looking at the camera — winks included, since a shut eye's
/// box keeps its width — sat at 0.78 or more. 0.6 splits the gap.
pub const HIDDEN_EYE_RATIO: f32 = 0.6;
/// What the reading says, for a screen or a filter.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum EyeState {
/// Every eye that could be read is open.
Open,
/// An eye that could be read is closed — a blink, or a wink.
Closed,
/// The eyes cannot be seen. Neither open nor closed, and a filter for
/// either leaves the face alone.
Sunglasses,
/// No eye was sharp enough, large enough and in view to read. Neither
/// open nor closed, like sunglasses, and left alone by every filter.
Unreadable,
}
impl Eye {
/// Whether this eye can be read at all: enough pixels, sharp enough,
/// and not the collapsed contour of a hidden eye — measured against
/// `other`, its partner.
pub fn readable(&self, other: &Eye) -> bool {
self.px >= MIN_EYE_PX
&& self.sharpness >= MIN_EYE_SHARPNESS
&& self.px >= other.px * HIDDEN_EYE_RATIO
}
}
impl EyeReading {
pub fn state(&self) -> EyeState {
if self.sunglasses >= SUNGLASSES_THRESHOLD {
return EyeState::Sunglasses;
}
let readable = [
self.right.readable(&self.left).then_some(self.right.open),
self.left.readable(&self.right).then_some(self.left.open),
];
let mut any = false;
for open in readable.into_iter().flatten() {
any = true;
if open < EYES_OPEN_THRESHOLD {
return EyeState::Closed;
}
}
if any {
EyeState::Open
} else {
EyeState::Unreadable
}
}
/// Whether this is a face a "no one blinking" filter should drop.
///
/// The filter's question, rather than [`EyeState`]'s four-way answer,
/// because the two differ on exactly the cases that matter: a face
/// behind sunglasses, or one whose eyes could not be read, is not open
/// — and it is not a blink either. Only [`EyeState::Closed`] is one.
pub fn is_blink(&self) -> bool {
self.state() == EyeState::Closed
}
}
impl EyeState {
/// The word the People screen puts on the face.
pub fn label(&self) -> &'static str {
match self {
EyeState::Open => "Eyes open",
EyeState::Closed => "Eyes closed",
EyeState::Sunglasses => "Sunglasses",
EyeState::Unreadable => "Eyes unclear",
}
}
}
#[cfg(test)]
mod tests {
use super::*;
fn eye(open: f32) -> Eye {
Eye {
open,
px: 40.0,
sharpness: 0.1,
}
}
fn reading(right: f32, left: f32, sunglasses: f32) -> EyeReading {
EyeReading {
right: eye(right),
left: eye(left),
sunglasses,
}
}
#[test]
fn both_eyes_open_is_open() {
assert_eq!(reading(0.9, 0.8, 0.1).state(), EyeState::Open);
assert!(!reading(0.9, 0.8, 0.1).is_blink());
}
/// A wink is not "eyes open": one eye closed lands the same place a
/// blink does, and a filter for "nobody blinking" should drop it.
#[test]
fn one_eye_closed_is_closed() {
assert_eq!(reading(0.9, 0.2, 0.1).state(), EyeState::Closed);
assert_eq!(reading(0.2, 0.9, 0.1).state(), EyeState::Closed);
assert!(reading(0.2, 0.9, 0.1).is_blink());
}
/// The whole reason the sunglasses number exists: whatever the eye
/// classifier says over dark glass, it is not a reading of the eyes.
#[test]
fn sunglasses_override_the_eye_readings_either_way() {
assert_eq!(reading(0.9, 0.9, 0.8).state(), EyeState::Sunglasses);
assert_eq!(reading(0.1, 0.1, 0.8).state(), EyeState::Sunglasses);
assert!(!reading(0.1, 0.1, 0.8).is_blink());
}
/// A soft or tiny eye is not asked; if neither can be, the face is
/// unreadable rather than closed.
#[test]
fn a_soft_or_tiny_eye_is_not_read() {
let mut r = reading(0.1, 0.9, 0.0);
r.right.sharpness = MIN_EYE_SHARPNESS / 2.0;
assert_eq!(r.state(), EyeState::Open, "the soft closed eye is ignored");
let mut r = reading(0.1, 0.9, 0.0);
r.right.px = MIN_EYE_PX - 1.0;
assert_eq!(r.state(), EyeState::Open, "the tiny closed eye is ignored");
let mut r = reading(0.1, 0.1, 0.0);
r.right.sharpness = 0.0;
r.left.px = 3.0;
assert_eq!(r.state(), EyeState::Unreadable);
assert!(!r.is_blink());
assert_eq!(r.state().label(), "Eyes unclear");
}
/// A profile: the far eye's contour collapses, and the sliver is not
/// read. The near eye still decides.
#[test]
fn a_turned_heads_collapsed_far_eye_is_not_read() {
let mut r = reading(0.05, 0.95, 0.0);
r.right.px = 40.0 * HIDDEN_EYE_RATIO - 1.0;
assert!(!r.right.readable(&r.left));
assert_eq!(r.state(), EyeState::Open);
let mut blink = reading(0.95, 0.05, 0.0);
blink.right.px = 40.0 * HIDDEN_EYE_RATIO - 1.0;
assert_eq!(blink.state(), EyeState::Closed);
// Both eyes narrow but alike is not a turned head: both count.
let mut small = reading(0.05, 0.95, 0.0);
small.right.px = 14.0;
small.left.px = 14.0;
assert_eq!(small.state(), EyeState::Closed);
}
#[test]
fn the_thresholds_are_inclusive_at_the_decision_point() {
assert_eq!(
reading(EYES_OPEN_THRESHOLD, EYES_OPEN_THRESHOLD, 0.0).state(),
EyeState::Open
);
assert_eq!(
reading(1.0, 1.0, SUNGLASSES_THRESHOLD).state(),
EyeState::Sunglasses
);
let mut r = reading(1.0, 1.0, 0.0);
r.right.px = MIN_EYE_PX;
r.left.px = MIN_EYE_PX;
r.right.sharpness = MIN_EYE_SHARPNESS;
assert!(r.right.readable(&r.left));
}
}
+263
View File
@@ -0,0 +1,263 @@
//! TRACES: FR-CULL-8a
//! Dense facial landmarks — InsightFace's `2d106det` (docs/faces.md §17.2).
//!
//! SCRFD's five points place a face; they do not place an eye. Its eye
//! point is loose enough that a window centred on it left the eye in a
//! corner on turned and smiling heads, and two model-free ways of
//! re-centring it made things worse. So a second model draws the eye's lid
//! contour, and the eye box is cut from that.
//!
//! **Why this one.** Three were measured on the same faces — MediaPipe Face
//! Mesh V2, PIPNet and this — and tied on what the eye classifier made of
//! their boxes (22 of 25 open eyes read open, against 19 from the SCRFD
//! point). This is the cheapest of the three by a wide margin (5 MB, 106
//! points, ~24 ms in tract), and it is under the grant the detector and
//! embedder already carry rather than a new one to read.
//!
//! # Pre-processing
//!
//! Ported from InsightFace's `landmark.py`: a square crop centred on the
//! detector box, 1.5× its longer edge, resized to 192; **RGB in 0..255**
//! (the graph carries its own `bn_data` normalisation, so `input_mean` is
//! 0 and `input_std` 1); 106 `(x, y)` in −1..1 mapped back through
//! `(p + 1) · 96`. The graph's batch dimension is the literal `None` and
//! is pinned to 1 by `tools/fix-face-model-shapes.sh`, like the embedder's.
//!
//! # The layout
//!
//! Checked by drawing the points on the reference faces rather than taken
//! from a diagram: the subject's right eye (image-left) is points 33–42,
//! the left 87–96, ten each round the lids.
use ndarray::Array4;
use crate::align::crop_box;
use crate::{FaceError, Pixels};
use dr_inference_engine::{Form, Model, Role};
/// The graph's input edge, in pixels.
pub const INPUT_EDGE: usize = 192;
/// How many points the model returns.
pub const POINTS: usize = 106;
/// The crop's edge as a multiple of the detector box's longer edge.
const CROP_SCALE: f32 = 1.5;
/// The span of the frame, in long-edge units, the packed form covers: a
/// quarter of the frame outside each edge.
pub const PACKED_RANGE: (f32, f32) = (-0.25, 1.25);
/// Bytes the packed form of one face's landmarks takes.
pub const PACKED_BYTES: usize = POINTS * 4;
/// Point indices of the subject's right eye's lid contour (image-left).
pub const RIGHT_EYE: [usize; 10] = [33, 34, 35, 36, 37, 38, 39, 40, 41, 42];
/// Point indices of the subject's left eye's lid contour (image-right).
pub const LEFT_EYE: [usize; 10] = [87, 88, 89, 90, 91, 92, 93, 94, 95, 96];
/// The 106 points of one face, in **source pixels**.
#[derive(Debug, Clone, PartialEq)]
pub struct Landmarks {
pub points: [(f32, f32); POINTS],
}
impl Landmarks {
/// Storage form: `106 × (x, y)` as little-endian **`u16` fixed point**
/// over the frame, 424 bytes.
///
/// Each coordinate is normalised by `long_edge` like the five points the
/// catalog already keeps, then mapped over [`PACKED_RANGE`] — a quarter
/// of the frame either side of it, because a landmark on a face at the
/// edge does land outside the image — onto 0..65535. That is 0.14 source
/// pixels on a 6000-pixel frame. `f16` would be the same size and worse:
/// its three significant figures near 1.0 are six pixels at that scale,
/// and the eye contour this is kept for is drawn to the pixel.
pub fn to_packed_bytes(&self, long_edge: f32) -> Vec<u8> {
let (lo, hi) = PACKED_RANGE;
let pack = |v: f32| -> [u8; 2] {
let t = ((v / long_edge - lo) / (hi - lo)).clamp(0.0, 1.0);
((t * 65535.0).round() as u16).to_le_bytes()
};
let mut out = Vec::with_capacity(POINTS * 4);
for &(x, y) in &self.points {
out.extend_from_slice(&pack(x));
out.extend_from_slice(&pack(y));
}
out
}
/// [`Self::to_packed_bytes`] read back, into source pixels of a frame
/// with this `long_edge`. `None` for a blob of the wrong length.
pub fn from_packed_bytes(bytes: &[u8], long_edge: f32) -> Option<Self> {
if bytes.len() != POINTS * 4 {
return None;
}
let (lo, hi) = PACKED_RANGE;
let unpack = |b: &[u8]| -> f32 {
let t = u16::from_le_bytes([b[0], b[1]]) as f32 / 65535.0;
(t * (hi - lo) + lo) * long_edge
};
let mut points = [(0.0_f32, 0.0_f32); POINTS];
for (i, p) in points.iter_mut().enumerate() {
let at = i * 4;
*p = (unpack(&bytes[at..at + 2]), unpack(&bytes[at + 2..at + 4]));
}
Some(Self { points })
}
/// The lid contour of the subject's right eye.
pub fn right_eye(&self) -> [(f32, f32); 10] {
RIGHT_EYE.map(|i| self.points[i])
}
/// The lid contour of the subject's left eye.
pub fn left_eye(&self) -> [(f32, f32); 10] {
LEFT_EYE.map(|i| self.points[i])
}
}
/// A loaded `2d106det` graph.
pub struct Landmarker {
session: Model,
}
impl Landmarker {
pub fn from_path(path: impl AsRef<std::path::Path>) -> Result<Self, FaceError> {
let bytes = std::fs::read(path).map_err(FaceError::ModelRead)?;
Self::from_bytes(&bytes)
}
pub fn from_bytes(bytes: &[u8]) -> Result<Self, FaceError> {
let model = dr_inference_engine::open(Role::Landmarks, Form::F32, bytes)?;
let acquired = model.acquire()?;
let session = acquired.lock();
let input = session.inputs().first().ok_or(FaceError::WrongModel {
expected: "2d106det",
detail: "model has no inputs".into(),
})?;
let shape: Option<Vec<i64>> = input.dtype().tensor_shape().map(|s| s.to_vec());
let want = [1, 3, INPUT_EDGE as i64, INPUT_EDGE as i64];
if shape.as_deref() != Some(&want[..]) {
return Err(FaceError::WrongModel {
expected: "2d106det",
detail: format!(
"input '{}' is {:?}, expected {:?} (batch pinned to 1)",
input.name(),
shape,
want
),
});
}
let out = session.outputs().first().ok_or(FaceError::WrongModel {
expected: "2d106det",
detail: "model has no outputs".into(),
})?;
let last: Option<i64> = out.dtype().tensor_shape().and_then(|d| d.last().copied());
if last != Some((POINTS * 2) as i64) {
return Err(FaceError::WrongModel {
expected: "2d106det",
detail: format!(
"output '{}' is {:?}-wide, expected {}",
out.name(),
last,
POINTS * 2
),
});
}
drop(session);
drop(acquired);
Ok(Self { session: model })
}
/// The landmarks of the face in `bbox` — `(x0, y0, x1, y1)` in source
/// pixels, the detector's box — read from the source.
///
/// `None` for a box with no area or a buffer that is not the size it
/// claims, as every crop here.
pub fn landmarks(
&mut self,
px: Pixels<'_>,
width: usize,
height: usize,
bbox: (f32, f32, f32, f32),
) -> Result<Option<Landmarks>, FaceError> {
let (w, h) = (bbox.2 - bbox.0, bbox.3 - bbox.1);
let side = w.max(h) * CROP_SCALE;
let (cx, cy) = ((bbox.0 + bbox.2) / 2.0, (bbox.1 + bbox.3) / 2.0);
let (x0, y0) = (cx - side / 2.0, cy - side / 2.0);
let Some(crop) = crop_box(
px,
width,
height,
(x0, y0, side, side),
INPUT_EDGE,
INPUT_EDGE,
) else {
return Ok(None);
};
let e = INPUT_EDGE;
let mut input = Array4::<f32>::zeros((1, 3, e, e));
for y in 0..e {
for x in 0..e {
for c in 0..3 {
input[[0, c, y, x]] = crop[(y * e + x) * 3 + c] * 255.0;
}
}
}
let acquired = self.session.acquire()?;
let mut session = acquired.lock();
let outputs = session
.run(ort::inputs![
ort::value::Tensor::from_array(input).map_err(FaceError::Inference)?
])
.map_err(FaceError::Inference)?;
let (_, data) = outputs[0]
.try_extract_tensor::<f32>()
.map_err(FaceError::Inference)?;
if data.len() < POINTS * 2 {
return Err(FaceError::WrongModel {
expected: "2d106det",
detail: format!("got {} values, expected {}", data.len(), POINTS * 2),
});
}
// −1..1 in the crop → crop pixels → source pixels.
let scale = side / e as f32;
let half = e as f32 / 2.0;
let mut points = [(0.0_f32, 0.0_f32); POINTS];
for (i, p) in points.iter_mut().enumerate() {
let (u, v) = ((data[2 * i] + 1.0) * half, (data[2 * i + 1] + 1.0) * half);
*p = (x0 + u * scale, y0 + v * scale);
}
Ok(Some(Landmarks { points }))
}
}
#[cfg(test)]
mod tests {
use super::*;
/// Packed and unpacked, every point comes back within a fifth of a
/// source pixel on a 6000-pixel frame — including one outside the
/// image, which a face at the edge does produce.
#[test]
fn dense_landmarks_round_trip_through_their_packed_bytes() {
let mut points = [(0.0_f32, 0.0_f32); POINTS];
for (i, p) in points.iter_mut().enumerate() {
*p = (i as f32 * 37.3 - 200.0, 5900.0 - i as f32 * 11.1);
}
let lm = Landmarks { points };
let bytes = lm.to_packed_bytes(6000.0);
assert_eq!(bytes.len(), PACKED_BYTES);
assert_eq!(PACKED_BYTES, 424);
let back = Landmarks::from_packed_bytes(&bytes, 6000.0).unwrap();
for (a, b) in lm.points.iter().zip(back.points.iter()) {
assert!((a.0 - b.0).abs() < 0.2, "{} vs {}", a.0, b.0);
assert!((a.1 - b.1).abs() < 0.2, "{} vs {}", a.1, b.1);
}
assert!(Landmarks::from_packed_bytes(&bytes[..100], 6000.0).is_none());
}
}
+59 -21
View File
@@ -1,8 +1,10 @@
//! Faces and identity (S14, docs/faces.md).
//!
//! Two models, run over the proxy tier, producing per face a box, five
//! Two models, run over the native render, producing per face a box, five
//! landmarks, a confidence and a 512-d embedding (FR-CULL-8) — and then the
//! arithmetic that turns embeddings into people (FR-CULL-9, FR-CULL-10).
//! arithmetic that turns embeddings into people (FR-CULL-9, FR-CULL-10). Two
//! more, optional, read each face's eyes and whether sunglasses hide them
//! (FR-CULL-8a, [`classify`] and [`eyes`]).
//!
//! Like `dr-segment`, this crate is **device-free**: no GPU adapter, no
//! Slint, nothing that needs a display. Unlike `dr-segment`, it carries **no
@@ -20,7 +22,8 @@
//! runtime; this crate takes bytes and never fetches anything.
//!
//! docs/faces.md §2 is the full reading, including what would have to change
//! for that to stop being true.
//! for that to stop being true. The eye-state models are the exception: MIT,
//! weights and all, and shipped in `models/face/` (docs/faces.md §17).
//!
//! # Why the runtime is split behind a feature
//!
@@ -34,26 +37,65 @@
pub mod align;
pub mod assign;
pub mod calibrate;
#[cfg(feature = "inference")]
pub mod classify;
pub mod cluster;
#[cfg(feature = "inference")]
pub mod detect;
#[cfg(feature = "inference")]
pub mod embed;
pub mod embedding;
pub mod eyes;
#[cfg(feature = "inference")]
pub mod landmarks;
pub mod naming;
pub mod neighbours;
pub use align::{warp, Aligned112, Similarity, ALIGNED_EDGE, ARCFACE_TEMPLATE};
/// Smallest long edge a face crop may be sampled from.
///
/// **A floor on the crop source, not on the detector input.** The distinction
/// is the whole of FR-CULL-8 and `docs/faces.md` §7: detection letterboxes
/// every buffer into 640×640, so its input resolution decides nothing, while
/// [`warp`] samples the 112×112 the embedder sees and so converts source
/// resolution directly into embedding quality. FR-CULL-8 requires that crop to
/// come from the native render; this is the guard that catches a caller
/// sampling from a proxy instead.
///
/// 1025 rather than 1024 because 1024 is exactly `dr_thumbs::ThumbSize::Large`,
/// the stored proxy tier a caller is most likely to reach for by mistake, so
/// the floor has to exclude it rather than admit it. Written as a minimum so
/// the test is `edge < MIN_CROP_EDGE` with no boundary to get wrong.
///
/// It is a coarse guard and deliberately so: whether any *individual* crop was
/// upsampled is answered exactly by `faces.crop_px` against [`ALIGNED_EDGE`],
/// and that is the number §7b measures. This only stops a whole pass reading
/// from the wrong tier.
pub const MIN_CROP_EDGE: u32 = 1025;
pub use align::{
crop_box, eye_box, eye_patch, head_views, warp, warp_pixels, Aligned112, EyePatch, HeadViews,
Pixels, Similarity, ALIGNED_EDGE, ARCFACE_TEMPLATE,
};
pub use assign::{identity_shares, RIVAL_FLOOR, TOP_MATCHES};
pub use calibrate::{Calibration, Pairs, ReliabilityBand};
#[cfg(feature = "inference")]
pub use classify::{EyeClassifier, EyeModels, SunglassesClassifier};
pub use cluster::{
cluster, cluster_scored, split, Candidate, Cluster, Grouping, DEFAULT_MERGE_PROBABILITY,
};
#[cfg(feature = "inference")]
pub use detect::{DetectOptions, Detection, Detector};
#[cfg(feature = "inference")]
pub use embed::Embedder;
pub use embedding::{Embedding, ModelId, EMBEDDING_DIM};
pub use embed::{Embedded, Embedder};
pub use embedding::{
in_gallery, read_f16_bytes, Embedding, ModelId, EMBEDDING_DIM, MIN_GALLERY_QUALITY,
};
pub use eyes::{
Eye, EyeReading, EyeState, EYES_OPEN_THRESHOLD, HIDDEN_EYE_RATIO, MIN_EYE_PX,
MIN_EYE_SHARPNESS, SUNGLASSES_THRESHOLD,
};
#[cfg(feature = "inference")]
pub use landmarks::{Landmarker, Landmarks};
pub use naming::{name_for_instance, name_instances, NamedFace};
/// What can go wrong between an image and a face.
@@ -82,25 +124,21 @@ pub enum FaceError {
ImageShape { expected: usize, got: usize },
}
/// Install tract as `ort`'s backend.
///
/// Idempotent, and it must happen before any other `ort` call: with
/// `alternative-backend` there is no linked runtime to fall back on, so an
/// un-set API is a panic rather than a slow path. Same helper as
/// `dr-segment::semantic`, for the same reason.
#[cfg(feature = "inference")]
pub(crate) fn install_backend() {
use std::sync::Once;
static ONCE: Once = Once::new();
ONCE.call_once(|| {
let _ = ort::set_api(ort_tract::api());
});
impl From<dr_inference_engine::Error> for FaceError {
fn from(e: dr_inference_engine::Error) -> Self {
match e {
dr_inference_engine::Error::Inference(e) => FaceError::Inference(e),
dr_inference_engine::Error::Io(e) => FaceError::ModelRead(e),
}
}
}
/// [`install_backend`] for the M1 probe example, which drives `ort` directly
/// rather than through [`detect::Detector`] so it can report the raw error.
/// Make sure `ort` has a backend, for the M1 probe example, which drives
/// `ort` directly rather than through [`detect::Detector`] so it can report
/// the raw error. Every other path goes through `dr-inference-engine`.
#[cfg(feature = "inference")]
#[doc(hidden)]
pub fn install_backend_for_probe() {
install_backend();
dr_inference_engine::ensure_runtime();
}
+47 -2
View File
@@ -117,6 +117,17 @@ pub struct Faces<'a> {
/// Which photograph each face came from. Two faces in one frame are not
/// the same person, so those pairs are never returned (docs/faces.md §9).
pub images: &'a [u64],
/// Which faces may be compared *against* — the gallery
/// ([`crate::embedding::MIN_GALLERY_QUALITY`]).
///
/// A pair needs at least one gallery side: a probe measured against a
/// reference is a comparison, two short vectors measured against each
/// other is noise agreeing with noise, and those pairs are never returned.
/// Filtered here rather than by the caller for the same reason
/// co-occurrence is: what this module leaves out of the list stays out of
/// the graph, the components and the merge order, so nothing downstream
/// has to remember the rule.
pub gallery: &'a [bool],
}
impl Faces<'_> {
@@ -272,8 +283,9 @@ fn scan_rows(scan: &Scan, from: usize, to: usize, out: &mut Vec<Pair>) {
let a = faces.row(i);
let crop_a = faces.crop_px[i];
let image_a = faces.images[i];
let gallery_a = faces.gallery[i];
for j in start..tile_end {
if image_a == faces.images[j] {
if image_a == faces.images[j] || !(gallery_a || faces.gallery[j]) {
continue;
}
let cos = dot(a, faces.row(j));
@@ -552,6 +564,7 @@ mod tests {
embeddings: Vec<f32>,
crop_px: Vec<f32>,
images: Vec<u64>,
gallery: Vec<bool>,
}
impl Set {
@@ -561,6 +574,7 @@ mod tests {
dim: DIM,
crop_px: &self.crop_px,
images: &self.images,
gallery: &self.gallery,
}
}
@@ -588,10 +602,12 @@ mod tests {
}
}
let crop_px = vec![150.0; embeddings.len()];
let gallery = vec![true; embeddings.len()];
Set {
embeddings: embeddings.concat(),
crop_px,
images,
gallery,
}
}
@@ -601,7 +617,7 @@ mod tests {
let mut out = Vec::new();
for i in 0..n {
for j in i + 1..n {
if faces.images[i] == faces.images[j] {
if faces.images[i] == faces.images[j] || !(faces.gallery[i] || faces.gallery[j]) {
continue;
}
let cos: f32 = faces
@@ -750,6 +766,35 @@ mod tests {
assert!(above_threshold(&s.faces(), &cal(), 0.9).is_empty());
}
/// A probe against a reference is a comparison; two probes against each
/// other is not. The rule lives here so that nothing downstream sees the
/// pair at all.
#[test]
fn two_faces_outside_the_gallery_are_never_paired() {
let mut s = population(1, 3, 1.0);
s.gallery = vec![false, false, true];
let pairs = above_threshold(&s.faces(), &cal(), 0.9);
assert!(
!pairs.iter().any(|p| p.i == 0 && p.j == 1),
"two probes were paired with each other"
);
// Each probe is still measured against the one reference.
assert!(pairs.iter().any(|p| p.i == 0 && p.j == 2));
assert!(pairs.iter().any(|p| p.i == 1 && p.j == 2));
}
#[test]
fn the_gallery_rule_matches_the_reference_at_scale() {
let mut s = population(60, 8, 0.97);
for (i, g) in s.gallery.iter_mut().enumerate() {
*g = i % 3 != 0;
}
let f = s.faces();
let got = above_threshold(&f, &cal(), 0.9);
let want = reference(&f, &cal(), 0.9);
assert!(same_pairs(&got, &want), "{} vs {}", got.len(), want.len());
}
#[test]
fn pairs_come_back_in_index_order() {
let s = population(300, 8, 0.97);
+5
View File
@@ -7,6 +7,11 @@ license.workspace = true
[dependencies]
dr-types.workspace = true
# The merge's geometry (FR-MRG-10): rotations, the focal length and the
# projections, solved on proxies by dr-pano and consumed here per chunk. The
# geometry alone — no keypoint model, no runtime — which is what the
# workspace entry turns off.
dr-pano.workspace = true
dr-decode.workspace = true
dr-pipeline.workspace = true
# The watershed's pixel passes are here because they are shaders; everything
+10 -9
View File
@@ -161,7 +161,7 @@ fn main() {
drain.set_param("saturation", ParamId("saturation"), -100.0);
// A touch of feather, or the colour stops dead on the model's outline and
// the eye goes straight to the edge instead of to the subject.
drain.feather = 0.02;
drain.base_mut().feather = 0.02;
pop.push(drain);
render_stack(
@@ -183,14 +183,14 @@ fn main() {
let mut brighter = subject_layer("m1", index, subject);
brighter.set_param("exposure", ParamId("exposure"), 0.45);
brighter.feather = 0.015;
brighter.base_mut().feather = 0.015;
lift.push(brighter);
let mut darker = subject_layer("m2", index, subject);
darker.invert = true;
darker.set_param("exposure", ParamId("exposure"), -0.55);
darker.set_param("saturation", ParamId("saturation"), -25.0);
darker.feather = 0.03;
darker.base_mut().feather = 0.03;
lift.push(darker);
render_stack(
@@ -221,9 +221,9 @@ fn main() {
let mut layer = subject_layer("m1", index, subject);
layer.invert = true;
layer.set_param("saturation", ParamId("saturation"), -100.0);
layer.feather = 0.004;
layer.morphology = morphology;
layer.morph_radius = radius;
layer.base_mut().feather = 0.004;
layer.base_mut().morphology = morphology;
layer.base_mut().morph_radius = radius;
stack.push(layer);
render_stack(
@@ -307,14 +307,14 @@ fn render_stack(
pw as usize,
ph as usize,
128,
match layer.morphology {
match layer.base().morphology {
Morphology::None => dr_segment::Morphology::None,
Morphology::Dilate => dr_segment::Morphology::Dilate,
Morphology::Erode => dr_segment::Morphology::Erode,
Morphology::Close => dr_segment::Morphology::Close,
Morphology::Open => dr_segment::Morphology::Open,
},
layer.morph_radius * pw.min(ph) as f32,
layer.base().morph_radius * pw.min(ph) as f32,
)
.distance
})
@@ -326,7 +326,7 @@ fn render_stack(
// after the framing map — which is what makes one mask correct at every
// output size, zoom and crop.
let array = masks
.render(stack, None, Some(&subjects), pw, ph)
.render(stack, None, Some(&subjects), Some(source), pw, ph)
.expect("rasterise masks");
let shader = compose_full(
@@ -335,6 +335,7 @@ fn render_stack(
ColourSpace::Srgb,
stack,
&SpotSet::new(),
&[],
);
adjust
.render_masked(source, &shader, ow, oh, Some(array))
+279
View File
@@ -0,0 +1,279 @@
//! Every way of editing a mask, as frames you can watch.
//!
//! The mask tools are hard to review from a still: what makes them right is
//! how the mask *moves* as a stroke is painted, as a correction is subtracted,
//! as an edge is shaped. This renders that — a synthetic photograph and one
//! frame per step of each mode — so the pipeline's behaviour can be watched
//! before any of it is wired to a finger.
//!
//! ```sh
//! cargo run -p dr-gpu --example mask_modes --release -- out
//! ffmpeg -y -framerate 12 -i out/paint-%03d.ppm out/paint.gif
//! ```
//!
//! PPM for the reason every other example here writes it: no encoder
//! dependency, and ffmpeg, ImageMagick and every viewer read it.
//!
//! # What it is really showing
//!
//! The fold that builds a layer's mask (`MaskPass::render`), through the
//! composed shader that samples it. A part drawn in the wrong order, a
//! subtraction that took the base with it, an erase stroke that punched
//! through the selection underneath — each of those is a frame here that looks
//! wrong, and none of them is visible in a single rendered still.
use dr_gpu::{AdjustPass, DemosaicedImage, GpuContext, MaskPass};
use dr_pipeline::descriptor::ParamId;
use dr_pipeline::mask::{Join, MaskLayer, MaskPart, MaskSource, MaskStack, Morphology};
use dr_pipeline::operation::compose_full;
use dr_pipeline::spot::SpotSet;
use dr_pipeline::{ops, Framing};
use dr_types::ColourSpace;
const W: u32 = 480;
const H: u32 = 320;
fn main() {
env_logger::init();
let dir = std::env::args().nth(1).unwrap_or_else(|| "out".into());
std::fs::create_dir_all(&dir).expect("output directory");
let Some(ctx) = pollster::block_on(GpuContext::new_headless()).ok() else {
eprintln!("no adapter; nothing to render");
return;
};
let source = DemosaicedImage::from_rgba8(&ctx, &scene(), W, H).expect("upload");
let mut masks = MaskPass::new(&ctx).expect("mask pass");
let mut adjust = AdjustPass::new(&ctx);
let mut shot = Shot {
source: &source,
masks: &mut masks,
adjust: &mut adjust,
dir: &dir,
};
write_ppm(&format!("{dir}/original.ppm"), &to_rgb(&scene()), W, H);
paint(&mut shot);
erase(&mut shot);
subtract(&mut shot);
invert(&mut shot);
shape(&mut shot);
println!("frames in {dir}/");
}
/// One layer, brightened hard, so the mask is legible rather than tasteful.
fn lifted(source: MaskSource) -> MaskLayer {
let mut layer = MaskLayer::new("m1", source);
layer.set_param("exposure", ParamId("exposure"), 1.6);
layer.set_param("saturation", ParamId("vibrance"), 0.6);
layer
}
/// A stroke painted from left to right across the subject, one frame per dab.
///
/// Each frame is the whole mask rasterised again, which is what the
/// application does today on every shape change — so the frames are also a
/// crude answer to "is a stroke's cost growing as it is painted".
fn paint(shot: &mut Shot) {
let mut layer = lifted(MaskSource::brush());
layer.begin_stroke(0, false, 0.13, 0.5, 1.0);
for (i, x) in steps(0.18, 0.82, 28).enumerate() {
layer.extend_stroke(0, x, 0.52 + 0.06 * (x * 9.0).sin());
shot.frame("paint", i, &layer);
}
layer.end_stroke(0);
}
/// The same layer, with an erase stroke taken back through the middle of it.
fn erase(shot: &mut Shot) {
let mut layer = lifted(MaskSource::brush());
layer.begin_stroke(0, false, 0.16, 0.5, 1.0);
for x in steps(0.18, 0.82, 20) {
layer.extend_stroke(0, x, 0.5);
}
layer.end_stroke(0);
for i in 0..8 {
shot.frame("erase", i, &layer);
}
layer.begin_stroke(0, true, 0.09, 0.7, 1.0);
for (i, x) in steps(0.25, 0.75, 20).enumerate() {
layer.extend_stroke(0, x, 0.5);
shot.frame("erase", 8 + i, &layer);
}
layer.end_stroke(0);
}
/// A correction joined to a radial selection and then taken out of it: the
/// part appears, is painted, and the mask loses exactly what it covers.
fn subtract(shot: &mut Shot) {
let mut layer = lifted(MaskSource::Radial {
centre: (0.5, 0.5),
radii: (0.42, 0.34),
angle: 0.0,
feather: 0.35,
});
for i in 0..8 {
shot.frame("subtract", i, &layer);
}
layer.push_part(MaskPart::painted("p2", Join::Subtract));
for (i, x) in steps(0.3, 0.72, 22).enumerate() {
if i == 0 {
layer.begin_stroke(1, false, 0.1, 0.6, 1.0);
}
layer.extend_stroke(1, x, 0.46);
shot.frame("subtract", 8 + i, &layer);
}
layer.end_stroke(1);
// And off again, which is the half a stroke cannot do: a part is a thing
// that can be switched off after the fact.
for i in 0..8 {
let mut without = layer.clone();
without.remove_part(1);
shot.frame("subtract", 30 + i, &without);
}
}
/// The layer turned over, and back, holding each state long enough to read.
fn invert(shot: &mut Shot) {
let mut layer = lifted(MaskSource::Radial {
centre: (0.42, 0.52),
radii: (0.3, 0.32),
angle: 0.0,
feather: 0.3,
});
for i in 0..24 {
layer.invert = (i / 8) % 2 == 1;
shot.frame("invert", i, &layer);
}
}
/// The edge controls, swept: a feather opening up, then a dilation pushing the
/// boundary out and an erosion pulling it back.
fn shape(shot: &mut Shot) {
let mut layer = lifted(MaskSource::Radial {
centre: (0.5, 0.5),
radii: (0.3, 0.3),
angle: 0.0,
feather: 0.02,
});
layer.base_mut().falloff = dr_pipeline::mask::Falloff::Smooth;
for (i, f) in steps(0.0, 0.09, 18).enumerate() {
layer.base_mut().feather = f;
shot.frame("shape", i, &layer);
}
layer.base_mut().morphology = Morphology::Dilate;
for (i, r) in steps(0.0, 0.06, 12).enumerate() {
layer.base_mut().morph_radius = r;
shot.frame("shape", 18 + i, &layer);
}
layer.base_mut().morphology = Morphology::Erode;
for (i, r) in steps(0.0, 0.06, 12).enumerate() {
layer.base_mut().morph_radius = r;
shot.frame("shape", 30 + i, &layer);
}
}
/// Everything one frame needs, so the mode functions read as what they do.
struct Shot<'a> {
source: &'a DemosaicedImage,
masks: &'a mut MaskPass,
adjust: &'a mut AdjustPass,
dir: &'a str,
}
impl Shot<'_> {
fn frame(&mut self, mode: &str, index: usize, layer: &MaskLayer) {
let mut stack = MaskStack::new();
stack.push(layer.clone());
let shader = compose_full(
&ops::chain(),
&Framing::new(),
ColourSpace::Srgb,
&stack,
&SpotSet::new(),
&[],
);
let array = self
.masks
.render(&stack, None, None, Some(self.source), W, H)
.expect("rasterise");
self.adjust
.render_masked(self.source, &shader, W, H, Some(array))
.expect("render");
let rgba = self.adjust.export_pixels().expect("readback").0;
write_ppm(
&format!("{}/{mode}-{index:03}.ppm", self.dir),
&to_rgb(&rgba),
W,
H,
);
}
}
/// `count` values from `from` to `to`, inclusive.
fn steps(from: f32, to: f32, count: usize) -> impl Iterator<Item = f32> {
(0..count).map(move |i| from + (to - from) * i as f32 / (count.max(2) - 1) as f32)
}
/// A picture with somewhere obvious to put a mask: a graded sky, a ground
/// band, and a warm subject sitting on the join.
fn scene() -> Vec<u8> {
let mut px = vec![0u8; (W * H * 4) as usize];
for y in 0..H {
for x in 0..W {
let (fx, fy) = (x as f32 / W as f32, y as f32 / H as f32);
let sky = [
(60.0 + 90.0 * fy) as u8,
(110.0 + 90.0 * fy) as u8,
(190.0 + 50.0 * fy) as u8,
];
let ground = [
(70.0 + 40.0 * fx) as u8,
(85.0 + 30.0 * fx) as u8,
(60.0 + 20.0 * fx) as u8,
];
let mut c = if fy > 0.62 { ground } else { sky };
// The subject: an ellipse, warm, with a little internal structure
// so a feathered edge has something to be soft against.
let (dx, dy) = ((fx - 0.5) / 0.22, (fy - 0.52) / 0.3);
if dx * dx + dy * dy < 1.0 {
let shade = 0.75 + 0.25 * (fx * 40.0).sin() * (fy * 30.0).cos();
c = [
(205.0 * shade) as u8,
(170.0 * shade) as u8,
(140.0 * shade) as u8,
];
}
let i = ((y * W + x) * 4) as usize;
px[i..i + 4].copy_from_slice(&[c[0], c[1], c[2], 255]);
}
}
px
}
fn to_rgb(rgba: &[u8]) -> Vec<u8> {
rgba.chunks_exact(4)
.flat_map(|p| [p[0], p[1], p[2]])
.collect()
}
fn write_ppm(path: &str, rgb: &[u8], w: u32, h: u32) {
use std::io::Write as _;
let mut f = std::io::BufWriter::new(std::fs::File::create(path).expect("create"));
write!(f, "P6\n{w} {h}\n255\n").expect("header");
f.write_all(rgb).expect("body");
}
+159
View File
@@ -0,0 +1,159 @@
//! Sweep a range mask's band and show what it selects.
//!
//! A diagnostic for FR-DEV-10. The other sweep example drives global
//! parameters through `EditGraph`; a band is not one of those — it lives on a
//! `MaskLayer`, is rasterised by its own pass, and only becomes visible
//! through whatever adjustment the layer carries.
//!
//! So the layer here is given a deliberately blunt adjustment — two stops down
//! — because the question this answers is *what does the band select*, not
//! *what would a photographer do with it*. A subtle edit would show a subtle
//! selection and prove nothing.
//!
//! ```sh
//! cargo run -p dr-gpu --release --example rangesweep -- IMG.CR2 out/ luminance 21
//! cargo run -p dr-gpu --release --example rangesweep -- IMG.CR2 out/ hue 21
//! ```
use dr_gpu::{AdjustPass, DemosaicedImage, Demosaicer, GpuContext, MaskPass};
use dr_pipeline::mask::{MaskLayer, MaskSource, MaskStack};
use dr_pipeline::operation::compose_full;
use dr_pipeline::ops;
use dr_pipeline::spot::SpotSet;
use dr_pipeline::{Framing, ParamId};
use dr_types::ColourSpace;
fn main() {
env_logger::init();
let mut args = std::env::args().skip(1);
let (Some(input), Some(out_dir), Some(mode)) = (args.next(), args.next(), args.next()) else {
eprintln!("usage: rangesweep <file.cr2> <out_dir> <luminance|hue|width> [steps]");
std::process::exit(2);
};
let steps: usize = args
.next()
.and_then(|s| s.parse().ok())
.filter(|n| *n >= 2)
.unwrap_or(21);
let ctx = pollster::block_on(GpuContext::new_headless()).expect("gpu");
let image = load(&ctx, &input);
std::fs::create_dir_all(&out_dir).expect("create out dir");
let (full_w, full_h) = image.size();
let longest = std::env::var("SWEEP_MAX_PX")
.ok()
.and_then(|s| s.parse::<u32>().ok())
.unwrap_or(1000);
let scale = (longest as f32 / full_w.max(full_h) as f32).min(1.0);
let w = ((full_w as f32 * scale) as u32).max(1);
let h = ((full_h as f32 * scale) as u32).max(1);
println!("rendering {w} x {h}");
let mut masks = MaskPass::new(&ctx).expect("mask pass");
let mut adjust = AdjustPass::new(&ctx);
for i in 0..steps {
let t = i as f32 / (steps - 1) as f32;
// What moves, and what the caption should say about it.
let (source, label) = match mode.as_str() {
// A half-wide band walking from black to white, so the selection
// sweeps across the tonal scale rather than merely widening.
"luminance" => {
let centre = t;
let half = 0.15;
(
MaskSource::luminance_range(centre - half, centre + half, 0.10),
format!("luminance band centred {:.2}", centre),
)
}
// The hue circle, at a fixed arc and a chroma floor that keeps the
// near-neutral parts of the picture out of it.
"hue" => (
MaskSource::colour_range(t, 0.08, 0.05, 1.0, 0.10),
format!("hue {:.0} deg", t * 360.0),
),
// The arc opening from nothing to everything, at a fixed hue.
"width" => (
MaskSource::colour_range(0.08, t * 0.5, 0.05, 1.0, 0.10),
format!("hue width {:.0} deg", t * 0.5 * 360.0),
),
other => {
eprintln!("unknown mode `{other}`");
std::process::exit(2);
}
};
let mut layer = MaskLayer::new("sweep", source);
// Two stops down: blunt on purpose. See the module docs.
layer.set_param("exposure", ParamId("exposure"), -2.0);
let mut stack = MaskStack::new();
stack.push(layer);
// No label field and no subject masks: a range needs neither. It is a
// weighting over the picture's own values, so the only input it wants
// is the picture, which is the `Some(&image)` below.
let array = masks
.render(&stack, None, None, Some(&image), w, h)
.expect("rasterise");
let shader = compose_full(
&ops::chain(),
&Framing::new(),
ColourSpace::Srgb,
&stack,
&SpotSet::new(),
&[],
);
adjust
.render_masked(&image, &shader, w, h, Some(array))
.expect("render");
let (pixels, pw, ph) = adjust.export_pixels().expect("readback");
let path = format!("{out_dir}/{mode}_{i:03}.ppm");
write_ppm(&path, &pixels, pw, ph);
std::fs::write(format!("{out_dir}/{mode}_{i:03}.txt"), format!("{label}\n"))
.expect("write label");
println!("{mode}[{i}] {label}");
}
}
fn write_ppm(path: &str, rgba: &[u8], w: u32, h: u32) {
use std::io::Write;
let mut out = Vec::with_capacity((w * h * 3) as usize + 32);
out.extend_from_slice(format!("P6\n{w} {h}\n255\n").as_bytes());
for px in rgba.chunks_exact(4) {
out.extend_from_slice(&px[..3]);
}
std::fs::File::create(path)
.expect("create output")
.write_all(&out)
.expect("write output");
}
/// Load either a RAW or an already-rendered image.
///
/// A JPEG takes the path `DemosaicedImage::from_rgba8` documents: no CFA to
/// interpolate, identity colour matrix, neutral white balance, and the shader
/// linearises the gamma-encoded pixels. The controls all still work; their
/// neutral is "as the camera left it" rather than "as the sensor recorded it",
/// which is worth knowing when reading a sweep made from one.
fn load(ctx: &GpuContext, path: &str) -> DemosaicedImage {
let bytes = std::fs::read(path).expect("read file");
match dr_decode::probe(&bytes) {
Some(dr_types::Format::Jpeg) => {
let p = dr_decode::decode_jpeg(&bytes).expect("decode jpeg");
println!("loaded {} x {} (rendered, not raw)", p.width, p.height);
DemosaicedImage::from_rgba8(ctx, &p.rgba, p.width, p.height).expect("upload")
}
_ => {
let raw = dr_decode::decode(&bytes).expect("decode raw");
println!("loaded {} x {} (raw)", raw.crop.width, raw.crop.height);
let demosaic = Demosaicer::new(ctx).expect("demosaicer");
demosaic.run(&raw).expect("demosaic")
}
}
}
+239
View File
@@ -0,0 +1,239 @@
//! Sweep one operation's parameters and write a frame per step.
//!
//! A diagnostic, not part of the product: it exists to show what a control
//! actually does to a photograph, one parameter at a time, so a new node can
//! be looked at rather than reasoned about.
//!
//! ```sh
//! cargo run -p dr-gpu --release --example sweep -- IMG.CR2 out/ colour_grading 21
//! ```
//!
//! It names no operation. The op id arrives as a string, the parameters and
//! their ranges come from the graph's own capabilities, and a node declared
//! yesterday sweeps on the same terms as one that shipped a year ago — which
//! is the property `ops/README.md` promises and the reason this is one example
//! rather than one per node.
//!
//! PPM out, like `develop.rs`, so it needs no encoder dependency; the caller
//! turns them into whatever it wants.
use dr_gpu::{AdjustPass, DemosaicedImage, Demosaicer, GpuContext};
use dr_pipeline::{Affects, EditGraph, OpId, ParamId, ParamKind};
use dr_types::ColourSpace;
fn main() {
env_logger::init();
let mut args = std::env::args().skip(1);
let (Some(input), Some(out_dir), Some(op)) = (args.next(), args.next(), args.next()) else {
eprintln!("usage: sweep <file.cr2> <out_dir> <op_id> [steps]");
eprintln!(" sweep <file.cr2> <out_dir> --list");
std::process::exit(2);
};
let steps: usize = args
.next()
.and_then(|s| s.parse().ok())
.filter(|n| *n >= 2)
.unwrap_or(21);
// Sweep one named parameter rather than all of them.
let only: Option<String> = args.next();
let ctx = pollster::block_on(GpuContext::new_headless()).expect("gpu");
let image = load(&ctx, &input);
let mut graph = EditGraph::default_chain();
// `--list` prints every op and parameter with its range, which is how the
// caller learns what there is to sweep without this file holding a list
// that would go stale.
if op == "--list" {
for cap in graph.capabilities() {
println!("{}", cap.id.0);
for p in &cap.params {
match p.kind {
ParamKind::Scalar { min, max, .. } => {
println!(" {:<20} {min} .. {max} default {}", p.id.0, p.default)
}
ParamKind::Bool => println!(" {:<20} bool", p.id.0),
ParamKind::Enum { ref variants } => {
println!(" {:<20} enum, {} variants", p.id.0, variants.len())
}
}
}
}
return;
}
std::fs::create_dir_all(&out_dir).expect("create out dir");
// `OpId`/`ParamId` hold `&'static str`, and an argument is not static.
// Leaking is right rather than expedient here: the ids live as long as the
// graph does, and this process exits immediately after.
let op_id = OpId(Box::leak(op.clone().into_boxed_str()));
let cap = graph
.capabilities()
.into_iter()
.find(|c| c.id == op_id)
.unwrap_or_else(|| {
eprintln!("no operation `{op}` — try --list");
std::process::exit(1);
});
// Bounded output rather than full sensor resolution. This is the proxy
// path FR-DSP-1 already renders through, so it is the same code the
// develop view uses — and a 25 MP frame would be a 75 MB PPM, times
// several hundred frames in one sweep.
let (full_w, full_h) = image.size();
let longest = std::env::var("SWEEP_MAX_PX")
.ok()
.and_then(|s| s.parse::<u32>().ok())
.unwrap_or(1100);
let scale = (longest as f32 / full_w.max(full_h) as f32).min(1.0);
let w = ((full_w as f32 * scale) as u32).max(1);
let h = ((full_h as f32 * scale) as u32).max(1);
println!("rendering {w} x {h} (from {full_w} x {full_h})");
let mut adjust = AdjustPass::new(&ctx);
// The reference frame: every parameter at its default. Written once so a
// viewer can see what the sweep is departing from.
render_to(
&mut adjust,
&image,
&graph,
w,
h,
&out_dir,
"neutral",
0,
0.0,
);
// Parameters held away from their default for the duration, as
// `SWEEP_HOLD=shadow_strength=70,midtone_hue=210`.
//
// Needed because a parameter is not always meaningful alone. Where a
// `presentation:` block groups several into one conceptual control — a
// hue and the strength behind it — sweeping one with the other at its
// default renders the same frame every time, which looks like a broken
// node rather than a correctly declared neutral.
let hold = std::env::var("SWEEP_HOLD").unwrap_or_default();
for clause in hold.split(',').filter(|c| !c.trim().is_empty()) {
let Some((name, value)) = clause.split_once('=') else {
eprintln!("SWEEP_HOLD wants name=value, got `{clause}`");
std::process::exit(2);
};
let value: f32 = value.trim().parse().expect("hold value");
let name = Box::leak(name.trim().to_string().into_boxed_str());
graph.set_param(op_id, ParamId(name), value);
println!("holding {name} = {value}");
}
for p in &cap.params {
if only.as_deref().is_some_and(|o| o != p.id.0) {
continue;
}
let ParamKind::Scalar { min, max, .. } = p.kind else {
eprintln!("skipping {} — only scalars sweep meaningfully", p.id.0);
continue;
};
let param_id = ParamId(Box::leak(p.id.0.to_string().into_boxed_str()));
for i in 0..steps {
let t = i as f32 / (steps - 1) as f32;
let value = min + (max - min) * t;
graph.set_param(op_id, param_id, value);
render_to(
&mut adjust,
&image,
&graph,
w,
h,
&out_dir,
p.id.0,
i,
value,
);
}
// Back to default before the next parameter, so each sweep is of one
// control rather than of everything tried so far.
graph.set_param(op_id, param_id, p.default);
}
}
#[allow(clippy::too_many_arguments)]
fn render_to(
adjust: &mut AdjustPass,
image: &dr_gpu::DemosaicedImage,
graph: &EditGraph,
w: u32,
h: u32,
dir: &str,
name: &str,
index: usize,
value: f32,
) {
// The detail path, always. A neighbourhood node — dehaze, clarity,
// sharpening — runs as its own dispatch after the fused pass, and the
// fused path refuses a shader composed with one rather than rendering it
// wrongly. `render_detailed` falls through to the plain path when the
// chain has no detail stage, so this one call serves both kinds of node
// and the example never has to know which it was handed.
let shader = graph.compose_for(ColourSpace::Srgb);
let scale = graph.render_scale(image.size(), (w, h));
let detail =
graph.compose_detail_for(scale.full_size(), scale.render_size(), ColourSpace::Srgb);
let key = graph.invalidation().through(Affects::Colour);
adjust
.render_detailed(image, &shader, w, h, None, &detail, key)
.expect("adjust");
let (pixels, pw, ph) = adjust.export_pixels().expect("readback");
let path = format!("{dir}/{name}_{index:03}.ppm");
write_ppm(&path, &pixels, pw, ph);
// The value goes beside the frame rather than into the filename: a caption
// wants "-37.5", and a filename that carried it would need escaping and
// would sort wrongly.
let meta = format!("{dir}/{name}_{index:03}.txt");
std::fs::write(meta, format!("{name} {value:.4}\n")).expect("write value");
println!("{name}[{index}] = {value:.4}");
}
fn write_ppm(path: &str, rgba: &[u8], w: u32, h: u32) {
use std::io::Write;
let mut out = Vec::with_capacity((w * h * 3) as usize + 32);
out.extend_from_slice(format!("P6\n{w} {h}\n255\n").as_bytes());
for px in rgba.chunks_exact(4) {
out.extend_from_slice(&px[..3]);
}
std::fs::File::create(path)
.expect("create output")
.write_all(&out)
.expect("write output");
}
/// Load either a RAW or an already-rendered image.
///
/// A JPEG takes the path `DemosaicedImage::from_rgba8` documents: no CFA to
/// interpolate, identity colour matrix, neutral white balance, and the shader
/// linearises the gamma-encoded pixels. The controls all still work; their
/// neutral is "as the camera left it" rather than "as the sensor recorded it",
/// which is worth knowing when reading a sweep made from one.
fn load(ctx: &GpuContext, path: &str) -> DemosaicedImage {
let bytes = std::fs::read(path).expect("read file");
match dr_decode::probe(&bytes) {
Some(dr_types::Format::Jpeg) => {
let p = dr_decode::decode_jpeg(&bytes).expect("decode jpeg");
println!("loaded {} x {} (rendered, not raw)", p.width, p.height);
DemosaicedImage::from_rgba8(ctx, &p.rgba, p.width, p.height).expect("upload")
}
_ => {
let raw = dr_decode::decode(&bytes).expect("decode raw");
println!("loaded {} x {} (raw)", raw.crop.width, raw.crop.height);
let demosaic = Demosaicer::new(ctx).expect("demosaicer");
demosaic.run(&raw).expect("demosaic")
}
}
}
+249 -9
View File
@@ -101,6 +101,16 @@ pub struct AdjustPass {
/// switched on does not build a pipeline layout mid-frame.
linear_bind_group_layout: wgpu::BindGroupLayout,
linear_pipeline_layout: wgpu::PipelineLayout,
/// TRACES: FR-MRG-2
/// A third layout, writing `rgba32float`, for the camera-space tap a
/// merge reads (`OutputMode::CameraLinear`). Same reasoning as the
/// linear one: the format is in the layout, so a format is a layout.
camera_bind_group_layout: wgpu::BindGroupLayout,
camera_pipeline_layout: wgpu::PipelineLayout,
/// The camera-space texture the last `render_camera_linear` wrote.
/// Separate from `targets`: a different format, and a merge reads it
/// back or samples it while the display targets go on being swapped.
camera_target: Option<Target>,
/// TRACES: FR-DEV-3d
/// What the linear intermediate currently holds, and at what size.
///
@@ -303,6 +313,11 @@ impl AdjustPass {
pub const FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::Rgba8Unorm;
/// TRACES: FR-MRG-2
/// The camera-space tap's format: full precision, because what it holds
/// is written back as a RAW at the sensor's own scale (FR-MRG-3).
pub const CAMERA_FORMAT: wgpu::TextureFormat = wgpu::TextureFormat::Rgba32Float;
pub fn new(ctx: &GpuContext) -> Self {
let bind_group_layout = Self::layout_writing(ctx, Self::FORMAT, "adjust-bgl");
@@ -330,6 +345,15 @@ impl AdjustPass {
bind_group_layouts: &[Some(&linear_bind_group_layout)],
immediate_size: 0,
});
let camera_bind_group_layout =
Self::layout_writing(ctx, Self::CAMERA_FORMAT, "adjust-camera-bgl");
let camera_pipeline_layout =
ctx.device
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
label: Some("adjust-camera-layout"),
bind_group_layouts: &[Some(&camera_bind_group_layout)],
immediate_size: 0,
});
// A 1x1 single-layer mask, bound when the edit has no local
// adjustments. The generated shader never samples it — no layer block
@@ -412,6 +436,9 @@ impl AdjustPass {
detail: DetailRunner::new(ctx),
linear_bind_group_layout,
linear_pipeline_layout,
camera_bind_group_layout,
camera_pipeline_layout,
camera_target: None,
colour_key: None,
colour_dispatches: 0,
detail_dispatches: 0,
@@ -549,6 +576,7 @@ impl AdjustPass {
let layout = match shader.output_mode {
OutputMode::Encoded => &self.pipeline_layout,
OutputMode::LinearWorking => &self.linear_pipeline_layout,
OutputMode::CameraLinear => &self.camera_pipeline_layout,
};
let pipeline =
@@ -1126,28 +1154,206 @@ impl AdjustPass {
self.copy_output()
}
/// TRACES: FR-MRG-2
/// Render the camera-space tap: the source after its lens warp and
/// nothing else, at full precision.
///
/// `shader` must come from `EditGraph::compose_camera_linear` — it is
/// refused otherwise, for the reason `render_masked` refuses a linear
/// one: the storage format is in the layout. The profile uniforms are
/// filled neutral here rather than from the source, which is the whole
/// point of the mode (`OutputMode::CameraLinear`): unit white balance,
/// identity matrix, base curve off. The non-linear flag is kept, so a
/// JPEG source is still linearised — camera space for a JPEG is the
/// decoded values made linear, which is the best that exists.
///
/// The texture stays on the device for a merge's warp to sample; see
/// [`Self::camera_texture`] and [`Self::read_camera_linear`].
pub fn render_camera_linear(
&mut self,
source: &DemosaicedImage,
shader: &ComposedShader,
width: u32,
height: u32,
) -> Result<&wgpu::Texture, GpuError> {
if shader.output_mode != OutputMode::CameraLinear {
return Err(GpuError::ShaderCompilation(
"render_camera_linear takes the shader from EditGraph::compose_camera_linear \
and no other; this one writes a different format"
.into(),
));
}
self.colour_key = None;
let (width, height) = (width.max(1), height.max(1));
self.ensure_camera_target(width, height);
let mut uniforms = Self::fused_uniforms(source, shader);
// Neutral profile: the numbers the sensor produced, and only those.
let non_linear = uniforms[15];
uniforms[0..4].copy_from_slice(&[1.0, 0.0, 0.0, 0.0]);
uniforms[4..8].copy_from_slice(&[0.0, 1.0, 0.0, 0.0]);
uniforms[8..12].copy_from_slice(&[0.0, 0.0, 1.0, 0.0]);
uniforms[12..16].copy_from_slice(&[1.0, 1.0, 1.0, non_linear]);
let b = dr_pipeline::BASE_CURVE_UNIFORM_OFFSET;
uniforms[b + 10] = 0.0;
let params_buf = self
.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("adjust-camera-params"),
contents: bytemuck::cast_slice(&uniforms),
usage: wgpu::BufferUsages::UNIFORM,
});
let _ = self.pipeline(shader)?;
let pipeline = self
.cache
.get(&shader.structure_hash)
.expect("compiled above");
let target = self.camera_target.as_ref().expect("ensured above");
let bind_group = self
.ctx
.device
.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("adjust-camera-bg"),
layout: &self.camera_bind_group_layout,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: wgpu::BindingResource::TextureView(source.view()),
},
wgpu::BindGroupEntry {
binding: 1,
resource: params_buf.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 2,
resource: wgpu::BindingResource::TextureView(&target.view),
},
wgpu::BindGroupEntry {
binding: 3,
resource: wgpu::BindingResource::TextureView(&self.empty_masks),
},
wgpu::BindGroupEntry {
binding: 4,
resource: wgpu::BindingResource::TextureView(self.film_curves_view()),
},
wgpu::BindGroupEntry {
binding: 5,
resource: wgpu::BindingResource::TextureView(self.film_lut_view()),
},
],
});
let mut enc = self
.ctx
.device
.create_command_encoder(&wgpu::CommandEncoderDescriptor {
label: Some("adjust-camera-encoder"),
});
{
let mut pass = enc.begin_compute_pass(&wgpu::ComputePassDescriptor {
label: Some("adjust-camera-pass"),
timestamp_writes: None,
});
pass.set_pipeline(pipeline);
pass.set_bind_group(0, &bind_group, &[]);
pass.dispatch_workgroups(width.div_ceil(8), height.div_ceil(8), 1);
}
self.ctx.queue.submit(Some(enc.finish()));
self.colour_dispatches += 1;
Ok(&self.camera_target.as_ref().expect("ensured above").texture)
}
/// The camera-space texture, if one has been rendered.
pub fn camera_texture(&self) -> Option<&wgpu::Texture> {
self.camera_target.as_ref().map(|t| &t.texture)
}
/// TRACES: FR-MRG-2
/// Read the camera-space tap back: tightly packed RGBA `f32`,
/// `width * height * 4` values, alpha 1.0 everywhere.
pub fn read_camera_linear(&self) -> Result<(Vec<f32>, u32, u32), GpuError> {
let Some(target) = self.camera_target.as_ref() else {
return Err(GpuError::Readback("no camera-space render yet".into()));
};
let (bytes, w, h) = Self::copy_texture(&self.ctx, &target.texture, w_h(target), 16)?;
let floats: Vec<f32> = bytes
.chunks_exact(4)
.map(|b| f32::from_le_bytes([b[0], b[1], b[2], b[3]]))
.collect();
Ok((floats, w, h))
}
fn ensure_camera_target(&mut self, width: u32, height: u32) {
if self
.camera_target
.as_ref()
.is_some_and(|t| t.width == width && t.height == height)
{
return;
}
let texture = self.ctx.device.create_texture(&wgpu::TextureDescriptor {
label: Some("adjust-camera-output"),
size: wgpu::Extent3d {
width,
height,
depth_or_array_layers: 1,
},
mip_level_count: 1,
sample_count: 1,
dimension: wgpu::TextureDimension::D2,
format: Self::CAMERA_FORMAT,
// Written by compute, sampled by a merge's warp, copied out for
// the CPU. Never handed to the compositor, so no RENDER_ATTACHMENT.
usage: wgpu::TextureUsages::STORAGE_BINDING
| wgpu::TextureUsages::TEXTURE_BINDING
| wgpu::TextureUsages::COPY_SRC,
view_formats: &[],
});
let view = texture.create_view(&Default::default());
self.camera_target = Some(Target {
texture,
view,
width,
height,
});
}
/// The transfer itself.
fn copy_output(&self) -> Result<(Vec<u8>, u32, u32), GpuError> {
let Some(target) = self.targets[self.current].as_ref() else {
return Err(GpuError::Readback("nothing rendered yet".into()));
};
let (w, h) = (target.width, target.height);
Self::copy_texture(&self.ctx, &target.texture, w_h(target), 4)
}
let unpadded = w * 4;
/// Copy a whole texture to the CPU, `bytes_per_pixel` wide, rows
/// unpadded. Shared by the display readback and the camera-space one.
fn copy_texture(
ctx: &GpuContext,
texture: &wgpu::Texture,
(w, h): (u32, u32),
bytes_per_pixel: u32,
) -> Result<(Vec<u8>, u32, u32), GpuError> {
let unpadded = w * bytes_per_pixel;
let align = wgpu::COPY_BYTES_PER_ROW_ALIGNMENT;
let padded = unpadded.div_ceil(align) * align;
let buf = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
let buf = ctx.device.create_buffer(&wgpu::BufferDescriptor {
label: Some("adjust-readback"),
size: (padded * h) as u64,
usage: wgpu::BufferUsages::COPY_DST | wgpu::BufferUsages::MAP_READ,
mapped_at_creation: false,
});
let mut enc = self.ctx.device.create_command_encoder(&Default::default());
let mut enc = ctx.device.create_command_encoder(&Default::default());
enc.copy_texture_to_buffer(
wgpu::TexelCopyTextureInfo {
texture: &target.texture,
texture,
mip_level: 0,
origin: wgpu::Origin3d::ZERO,
aspect: wgpu::TextureAspect::All,
@@ -1166,7 +1372,7 @@ impl AdjustPass {
depth_or_array_layers: 1,
},
);
self.ctx.queue.submit(Some(enc.finish()));
ctx.queue.submit(Some(enc.finish()));
let slice = buf.slice(..);
let (tx, rx) = std::sync::mpsc::channel();
@@ -1177,7 +1383,7 @@ impl AdjustPass {
// Polled rather than parked, and bounded rather than spun forever —
// see `readback::await_mapping`, which the histogram's own transfer
// shares for exactly the same reasons.
await_mapping(&self.ctx, &rx)?;
await_mapping(ctx, &rx)?;
let data = slice.get_mapped_range();
let mut out = Vec::with_capacity((unpadded * h) as usize);
@@ -1191,6 +1397,10 @@ impl AdjustPass {
}
}
fn w_h(t: &Target) -> (u32, u32) {
(t.width, t.height)
}
/// Number the lines of generated source, so a compiler error can be located.
pub(crate) fn numbered(src: &str) -> String {
src.lines()
@@ -1242,6 +1452,10 @@ mod tests {
// rather than about a camera's colour response.
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
@@ -1442,6 +1656,10 @@ mod tests {
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
@@ -1845,7 +2063,21 @@ mod tests {
fused_blocks += usize::from(point);
}
// The count the loop above accumulated, plus framing — which emits a
// The lens corrections, which are the third kind of block. They are
// not in `descriptors` — they rewrite coordinates rather than
// transform a colour, so they are not operations — and they run ahead
// of the fetch rather than in either stage the loop above sorts into.
let mut warp_blocks = 0;
for desc in g.warp_descriptors() {
let id = desc.id.0;
assert!(
shader.source.contains(&format!("---- warp: {id} ----")),
"{id} was armed above and did not reach the shader"
);
warp_blocks += 1;
}
// The counts the loops above accumulated, plus framing — which emits a
// stage of its own rather than an operation block and is not in
// `descriptors`. Asserted as well as the per-operation exclusive-or
// because the two catch different faults: the XOR catches an operation
@@ -1853,7 +2085,7 @@ mod tests {
// in the chain asked for.
assert_eq!(
shader.source.matches("---- ").count(),
fused_blocks + 1,
fused_blocks + warp_blocks + 1,
"the fused shader carries a block nothing in the chain asked for"
);
assert!(
@@ -1955,6 +2187,10 @@ mod tests {
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
@@ -2055,6 +2291,10 @@ mod tests {
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
+232
View File
@@ -250,6 +250,135 @@ impl DemosaicedImage {
}
}
impl DemosaicedImage {
/// TRACES: FR-MRG-3
/// A source that is already RGB in camera space: a linear DNG, which is
/// what a merge writes. No demosaic; the samples are normalised by the
/// file's black and white levels exactly as the demosaic kernel would
/// normalise a photosite, and everything else — the matrix, the
/// balance, the body's base curve — is carried through as for a CFA
/// file, because the composite is developed as one photograph from the
/// body that took its sources.
pub fn from_linear_rgb16(ctx: &GpuContext, raw: &RawImage) -> Result<Self, GpuError> {
let (width, height) = (raw.crop.width.max(1), raw.crop.height.max(1));
let limits = ctx.device.limits();
if width > limits.max_texture_dimension_2d || height > limits.max_texture_dimension_2d {
return Err(GpuError::TooLarge(format!(
"{width}×{height} exceeds the device limit of {}",
limits.max_texture_dimension_2d
)));
}
let stride = raw.width as usize * 3;
let expected = raw.height as usize * stride;
if raw.data.len() < expected {
return Err(GpuError::TooLarge(format!(
"{} samples is short of the {expected} a {}×{} RGB image needs",
raw.data.len(),
raw.width,
raw.height
)));
}
let black = black_per_cell(raw);
let inv = inv_range_per_cell(raw);
// Per channel rather than per CFA cell: R, G, B are the first three.
let mut half: Vec<u16> = Vec::with_capacity((width * height * 4) as usize);
for y in 0..height as usize {
let row = (raw.crop.y as usize + y) * stride + raw.crop.x as usize * 3;
for x in 0..width as usize {
let p = &raw.data[row + x * 3..row + x * 3 + 3];
for c in 0..3 {
let v = (f32::from(p[c]) - black[c]) * inv[c];
half.push(f32_to_f16_bits_unclamped(v));
}
half.push(f32_to_f16_bits(1.0));
}
}
let texture = ctx.device.create_texture_with_data(
&ctx.queue,
&wgpu::TextureDescriptor {
label: Some("linear-rgb-source"),
size: wgpu::Extent3d {
width,
height,
depth_or_array_layers: 1,
},
mip_level_count: 1,
sample_count: 1,
dimension: wgpu::TextureDimension::D2,
format: Self::FORMAT,
usage: wgpu::TextureUsages::TEXTURE_BINDING | wgpu::TextureUsages::COPY_SRC,
view_formats: &[],
},
wgpu::util::TextureDataOrder::LayerMajor,
bytemuck::cast_slice(&half),
);
let view = texture.create_view(&Default::default());
Ok(Self {
texture,
view,
width,
height,
color_matrix: raw.color_matrix.unwrap_or(IDENTITY_3X3),
as_shot_wb: [raw.wb_coeffs[0], raw.wb_coeffs[1], raw.wb_coeffs[2]],
base_curve: raw.base_curve,
non_linear: false,
})
}
}
/// Convert an f32 to half-precision bits, the general case: sign,
/// subnormals, round-to-nearest-even, saturation at the largest finite.
///
/// `f32_to_f16_bits` below is the 8-bit special case and says why it can
/// be; this one exists because a linear DNG is not that case. A 14-bit
/// sensor's least significant step, normalised, is 6.1e-5 — right at f16's
/// smallest normal (6.1e-5) — so the deepest shadows of a composite land
/// in the subnormal range, and rounding them to zero would crush the
/// shadows of exactly the file that was written to keep them. Values below
/// zero (black subtraction on a noisy photosite) and above one (a highlight
/// past the white level) are legitimate and kept.
fn f32_to_f16_bits_unclamped(v: f32) -> u16 {
let bits = v.to_bits();
let sign = ((bits >> 16) & 0x8000) as u16;
let exp = ((bits >> 23) & 0xFF) as i32;
let mant = bits & 0x7F_FFFF;
if exp == 0xFF {
// Infinity or NaN: a NaN sample is a decode fault; store the largest
// finite rather than propagate it through a blend.
return sign | 0x7BFF;
}
let e = exp - 127 + 15;
if e >= 0x1F {
return sign | 0x7BFF;
}
if e <= 0 {
// Subnormal in f16 (or underflow). Shift the full mantissa with its
// implicit bit right by the deficit, rounding to nearest even.
if e < -10 {
return sign;
}
let m = (mant | 0x80_0000) >> (1 - e);
let shift = 13;
let rounded = round_shift(m, shift);
return sign | rounded as u16;
}
let rounded = round_shift(mant, 13);
// Rounding can carry into the exponent; that is correct.
sign | (((e as u32) << 10) + rounded) as u16
}
/// `v >> shift`, rounded to nearest with ties to even.
fn round_shift(v: u32, shift: u32) -> u32 {
let half = 1u32 << (shift - 1);
let mask = (1u32 << shift) - 1;
let low = v & mask;
let mut out = v >> shift;
if low > half || (low == half && (out & 1) == 1) {
out += 1;
}
out
}
/// Convert an f32 to IEEE 754 half-precision bits.
///
/// Written out rather than pulled in as a dependency: the inputs here are
@@ -389,6 +518,9 @@ impl Demosaicer {
/// `RawImage`; which of the two CFA families it came off is this
/// function's problem, not theirs.
pub fn run(&self, raw: &RawImage) -> Result<DemosaicedImage, GpuError> {
if raw.samples_per_pixel == 3 {
return DemosaicedImage::from_linear_rgb16(&self.ctx, raw);
}
let (width, height) = (raw.crop.width.max(1), raw.crop.height.max(1));
let limits = self.ctx.device.limits();
@@ -827,6 +959,43 @@ mod tests {
use super::*;
use dr_decode::CropRect;
fn f16_to_f32(bits: u16) -> f32 {
let sign = if bits & 0x8000 != 0 { -1.0 } else { 1.0 };
let e = ((bits >> 10) & 0x1F) as i32;
let m = (bits & 0x3FF) as f32;
if e == 0 {
sign * m * 2f32.powi(-24)
} else {
sign * (1.0 + m / 1024.0) * 2f32.powi(e - 15)
}
}
#[test]
fn unclamped_half_keeps_shadows_signs_and_highlights() {
// A 14-bit LSB, normalised: subnormal in f16, and must not be zero.
let lsb = 1.0 / 16383.0;
let back = f16_to_f32(f32_to_f16_bits_unclamped(lsb));
assert!((back - lsb).abs() / lsb < 0.01, "{back} vs {lsb}");
// A quarter of that, still representable.
let tiny = lsb / 4.0;
let back = f16_to_f32(f32_to_f16_bits_unclamped(tiny));
assert!((back - tiny).abs() / tiny < 0.05, "{back} vs {tiny}");
// Below zero and above one survive.
assert!((f16_to_f32(f32_to_f16_bits_unclamped(-0.01)) + 0.01).abs() < 1e-5);
assert!((f16_to_f32(f32_to_f16_bits_unclamped(1.75)) - 1.75).abs() < 1e-3);
// Exact values are exact.
assert_eq!(f32_to_f16_bits_unclamped(1.0), 0x3C00);
assert_eq!(f32_to_f16_bits_unclamped(0.5), 0x3800);
assert_eq!(f32_to_f16_bits_unclamped(0.0), 0);
// Within one ULP of the clamped one on its domain: that one
// truncates the mantissa, this one rounds it.
for i in 0..=255 {
let v = i as f32 / 255.0;
let (a, b) = (f32_to_f16_bits_unclamped(v), f32_to_f16_bits(v));
assert!(a.abs_diff(b) <= 1, "{v}: {a} vs {b}");
}
}
fn raw_for(black: [u16; 4], white: u16) -> RawImage {
RawImage {
width: 4,
@@ -838,6 +1007,10 @@ mod tests {
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: None,
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
@@ -950,6 +1123,10 @@ mod tests {
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: None,
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
@@ -1170,6 +1347,53 @@ mod tests {
}
}
#[test]
fn a_grey_step_edge_stays_grey() {
// A flat patch cannot tell the Malvar kernels from any other set of
// weights that sum to zero. An edge can. A grey vertical step, so
// every photosite records the same profile, must come back with the
// three channels close together on both sides; any spread is false
// colour from interpolating across the edge.
//
// The bound is set by the paper's kernels, which peak at 0.19 here.
// With the ±2 terms of the green-site kernels transposed — the bug
// this test was written against — the peak is 0.375.
let Some(ctx) = ctx() else { return };
let d = Demosaicer::new(&ctx).expect("demosaicer");
let size = 32u32;
let white = 16383u16;
let mut raw = flat_cfa(CfaPattern::Rggb, size, [0, 0, 0], 0, white);
for y in 0..size {
for x in size / 2..size {
raw.data[(y * size + x) as usize] = white;
}
}
let img = d.run(&raw).expect("demosaic");
let px = read_rgba(&ctx, &img);
let (w, _) = img.size();
let mut worst = (0.0f32, 0u32, 0u32);
for y in 2..size - 2 {
for x in 2..size - 2 {
let p = px[(y * w + x) as usize];
let spread = (p[0] - p[1]).abs().max((p[2] - p[1]).abs());
if spread > worst.0 {
worst = (spread, x, y);
}
}
}
assert!(
worst.0 < 0.25,
"false colour of {} at ({}, {}) on a grey edge — the green-site \
kernels are interpolating across the edge",
worst.0,
worst.1,
worst.2
);
}
#[test]
fn output_is_free_of_nan_and_negatives() {
// f16 NaN propagates silently through every later stage; a negative
@@ -1196,6 +1420,10 @@ mod tests {
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: None,
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
@@ -1277,6 +1505,10 @@ mod tests {
],
color_matrix: None,
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
+17
View File
@@ -28,6 +28,7 @@ mod error;
mod focus;
mod histogram;
mod mask;
mod merge;
mod raw_histogram;
mod readback;
mod segment;
@@ -40,6 +41,7 @@ pub use demosaic::{DemosaicedImage, Demosaicer};
pub use detail::INTERMEDIATE_FORMAT as DETAIL_INTERMEDIATE_FORMAT;
pub use error::GpuError;
pub use focus::{FocusPeakPass, FocusPeaking, PeakColour, PeakSensitivity};
pub use merge::{Band, MergeFrame, MergeOutput, MergePass};
// Renamed on the way out: `BINS` says enough inside `histogram`, and nothing
// at all at a crate root shared with demosaic and segmentation.
pub use histogram::{Histogram, HistogramPass, BINS as HISTOGRAM_BINS};
@@ -281,6 +283,7 @@ impl GpuContext {
)))
}
/// TRACES: NFR-COMPAT-1
/// Ask one adapter for a device, with the limits the pipeline needs.
async fn device_from(adapter: &wgpu::Adapter) -> Result<(wgpu::Device, wgpu::Queue), GpuError> {
adapter
@@ -332,6 +335,20 @@ impl GpuContext {
pub fn backend(&self) -> wgpu::Backend {
self.adapter_info.backend
}
/// TRACES: NFR-OPS-1
/// The driver, as the adapter reported it, for a diagnostics bundle.
/// Name and version in one string because wgpu splits them by backend
/// and neither half means much without the other.
pub fn driver(&self) -> String {
let info = &self.adapter_info;
match (info.driver.is_empty(), info.driver_info.is_empty()) {
(true, true) => "unknown driver".to_string(),
(false, true) => info.driver.clone(),
(true, false) => info.driver_info.clone(),
(false, false) => format!("{} {}", info.driver, info.driver_info),
}
}
}
#[repr(C)]
+638 -98
View File
@@ -21,6 +21,17 @@
//! boxes, one draw each, compositing onto the slice with blend state — see the
//! second half of `mask.wgsl`.
//!
//! # The photograph, bound as an input
//!
//! A range mask (FR-DEV-10) selects by what a pixel *is*, so this pass reads
//! the demosaiced source as well as writing masks. It is bound for every draw
//! and looked at by two modes; everything else gets a 1x1 placeholder, for the
//! reason the label field below does — the bindings are fixed, and a second
//! pipeline differing only in what it ignores costs more than a texel.
//!
//! Nothing is read back and nothing is rasterised on this side. What crosses
//! into CPU memory for a range layer is five floats and a matrix.
//!
//! # The label field
//!
//! Region masks index a compacted label field uploaded once per segmentation.
@@ -30,10 +41,11 @@
//! `region_count`. The compaction is CPU-side and once per image, which is the
//! same place and cadence the region adjacency graph is already built at.
use dr_pipeline::mask::{MaskSource, MaskStack, Stroke, MAX_LAYERS};
use dr_pipeline::mask::{Join, MaskSource, MaskStack, Stroke, MAX_LAYERS};
use wgpu::util::DeviceExt;
use crate::{GpuContext, GpuError};
use crate::{DemosaicedImage, GpuContext, GpuError};
/// Modes understood by `mask.wgsl`. Kept beside the shader's `switch`.
const MODE_REGIONS: u32 = 0;
@@ -43,6 +55,10 @@ const MODE_SUBJECT: u32 = 3;
/// Brush layers go through their own entry points rather than the `switch`, so
/// this is only ever read by a person looking at a captured frame.
const MODE_BRUSH: u32 = 4;
/// TRACES: FR-DEV-10
const MODE_LUMINANCE: u32 = 5;
/// TRACES: FR-DEV-10
const MODE_COLOUR: u32 = 6;
/// Six vertices — two triangles — per stroke. See `vs_brush`.
const VERTICES_PER_STROKE: u32 = 6;
@@ -66,7 +82,25 @@ struct MaskParams {
axis: [f32; 2],
softness: f32,
angle: f32,
_pad1: [f32; 2],
/// TRACES: FR-DEV-10
/// Source texels per mask texel, per axis. See `image_value` in the
/// shader for why a range averages its footprint rather than sampling it.
source_step: [f32; 2],
/// Camera RGB → linear sRGB, one row per `vec4` because that is the
/// alignment a uniform gives a three-component vector anyway. Only a
/// range mask reads them.
cam_to_srgb: [[f32; 4]; 3],
/// `rgb`: as-shot white balance. `w`: non-zero for a gamma-encoded source.
/// The same packing the generated adjust shader uses, so the two agree by
/// construction rather than by inspection.
as_shot_wb: [f32; 4],
/// Whether this part is turned over before it joins the mask. Read by the
/// combine pass and by nothing else — see `fs_combine`.
invert: u32,
/// A uniform buffer is a multiple of sixteen bytes, and the flag above
/// takes four of them.
_pad: [u32; 3],
}
/// One stroke, as `mask.wgsl`'s `StrokeHeader` expects it.
@@ -355,6 +389,19 @@ pub struct MaskPass {
/// `dst(1 - a)` to erase.
brush_add: wgpu::RenderPipeline,
brush_erase: wgpu::RenderPipeline,
/// Reads a part back out of [`Self::scratch`] and blends it into the
/// layer's slice. The set operation is the blend state, so these two are
/// one shader as well.
combine_layout: wgpu::BindGroupLayout,
combine_union: wgpu::RenderPipeline,
combine_subtract: wgpu::RenderPipeline,
/// Where a part is drawn before it is joined.
///
/// One texture for the whole stack rather than one per layer, because
/// layers rasterise in sequence and a part is read back immediately after
/// it is drawn. Allocated the first time a layer has more than one part,
/// so a library of unedited masks never pays for it.
scratch: Option<Scratch>,
array: Option<MaskArray>,
/// How many times the array texture has been (re)allocated.
///
@@ -371,6 +418,14 @@ pub struct MaskPass {
/// label slots even when rasterising a gradient. A placeholder is cheaper
/// and far simpler than two pipelines differing only in what they ignore.
placeholder: LabelField,
/// TRACES: FR-DEV-10
/// Bound at the image slot for every mask that is not a range.
///
/// Never sampled by those modes, so its contents do not matter — but it is
/// cleared rather than left undefined, because a placeholder whose value
/// is arbitrary is one that makes a binding mistake look like a mask that
/// nearly works.
empty_image: wgpu::TextureView,
}
impl MaskPass {
@@ -405,6 +460,20 @@ impl MaskPass {
},
count: None,
},
// TRACES: FR-DEV-10
// The photograph, for a range mask. Unfilterable for the
// same reason the field above is: every read is a
// `textureLoad`, and this pipeline binds no sampler.
wgpu::BindGroupLayoutEntry {
binding: 6,
visibility: wgpu::ShaderStages::FRAGMENT,
ty: wgpu::BindingType::Texture {
sample_type: wgpu::TextureSampleType::Float { filterable: false },
view_dimension: wgpu::TextureViewDimension::D2,
multisampled: false,
},
count: None,
},
],
});
@@ -501,6 +570,88 @@ impl MaskPass {
"mask-brush-add",
blend_state(wgpu::BlendFactor::One, wgpu::BlendFactor::OneMinusSrc),
);
// The pipelines that join one part to the mask so far. The blend
// state is the set operation and the shader is the same three
// vertices either way — which is why adding a way to combine masks
// cost no shader arithmetic at all.
let combine_layout =
ctx.device
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
label: Some("mask-combine-bgl"),
entries: &[
uniform_entry(0),
wgpu::BindGroupLayoutEntry {
binding: 7,
visibility: wgpu::ShaderStages::FRAGMENT,
ty: wgpu::BindingType::Texture {
// Loaded texel by texel at matching size, so
// there is nothing to filter and no sampler.
sample_type: wgpu::TextureSampleType::Float { filterable: false },
view_dimension: wgpu::TextureViewDimension::D2,
multisampled: false,
},
count: None,
},
],
});
let combine_pipeline_layout =
ctx.device
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
label: Some("mask-combine-layout"),
bind_group_layouts: &[Some(&combine_layout)],
immediate_size: 0,
});
let combine = |label, blend| {
ctx.device
.create_render_pipeline(&wgpu::RenderPipelineDescriptor {
label: Some(label),
layout: Some(&combine_pipeline_layout),
vertex: wgpu::VertexState {
module: &module,
entry_point: Some("vs"),
compilation_options: Default::default(),
buffers: &[],
},
fragment: Some(wgpu::FragmentState {
module: &module,
entry_point: Some("fs_combine"),
compilation_options: Default::default(),
targets: &[Some(wgpu::ColorTargetState {
format: MaskArray::FORMAT,
blend: Some(blend),
write_mask: wgpu::ColorWrites::ALL,
})],
}),
primitive: wgpu::PrimitiveState::default(),
depth_stencil: None,
multisample: wgpu::MultisampleState::default(),
multiview_mask: None,
cache: None,
})
};
// `max`, not source-over: a union must not build up where two parts
// overlap. Two selections that both half-cover a pixel select it half
// — adding them would make the overlap of two soft edges harder than
// either, which is a seam exactly where a photographer joined two
// things to avoid one.
let combine_union = combine(
"mask-combine-union",
wgpu::BlendState {
color: MAX_BLEND,
alpha: MAX_BLEND,
},
);
// `dst * (1 - src)`, which is the erase blend one level up: what the
// mask had, minus what this part covers, in proportion to how much of
// it the part covers.
let combine_subtract = combine(
"mask-combine-subtract",
blend_state(wgpu::BlendFactor::Zero, wgpu::BlendFactor::OneMinusSrc),
);
// The same, with the deposit thrown away: coverage is only ever taken
// off what earlier strokes on this layer put down. There is no negative
// coverage to accumulate, so erasing an unpainted layer is a no-op
@@ -515,6 +666,7 @@ impl MaskPass {
}
let placeholder = LabelField::upload(ctx, &[0], 1, 1, 0)?;
let empty_image = empty_image(ctx);
// Everywhere outside, so a layer that somehow reaches this masks
// nothing rather than everything.
let empty_subject = SubjectMasks::upload(ctx, &[&[-1.0f32][..]], 1, 1)?;
@@ -526,10 +678,15 @@ impl MaskPass {
brush_layout,
brush_add,
brush_erase,
combine_layout,
combine_union,
combine_subtract,
scratch: None,
array: None,
allocations: 0,
placeholder,
empty_subject,
empty_image,
})
}
@@ -538,17 +695,49 @@ impl MaskPass {
/// `labels` may be `None` when no layer is a region mask; a region layer
/// without one is skipped rather than drawn wrong, since a mask that
/// silently covers the whole frame would apply an edit everywhere.
///
/// `source` is the photograph a range layer measures (FR-DEV-10), and it
/// is skipped on the same rule for the same reason: without it the shader
/// would read a blank placeholder, and a band that happens to contain
/// black would then cover the whole frame.
pub fn render(
&mut self,
stack: &MaskStack,
labels: Option<&LabelField>,
subjects: Option<&SubjectMasks>,
source: Option<&DemosaicedImage>,
width: u32,
height: u32,
) -> Result<&MaskArray, GpuError> {
self.render_revealing(stack, labels, subjects, source, width, height, None)
}
/// TRACES: FR-DEV-19c
/// [`Self::render`], also drawing the layer being looked at.
///
/// A selection with no adjustment on it changes no pixel, so it is not
/// active and has no slice — which is right until somebody asks to *see*
/// it, and that is the state a photographer is in from choosing a subject
/// until deciding what to do to it.
///
/// `reveal` has to be the same one the shader was composed with and the
/// same one the distance fields were built for: all three index this array
/// by position in [`MaskStack::rendered`], and two of them disagreeing
/// shows as an adjustment applied through another layer's mask.
#[allow(clippy::too_many_arguments)]
pub fn render_revealing(
&mut self,
stack: &MaskStack,
labels: Option<&LabelField>,
subjects: Option<&SubjectMasks>,
source: Option<&DemosaicedImage>,
width: u32,
height: u32,
reveal: Option<&dr_pipeline::mask::Reveal>,
) -> Result<&MaskArray, GpuError> {
// At least one layer, because a zero-layer texture array is invalid
// and the shader binds this slot unconditionally.
let active = stack.active_count().clamp(1, MAX_LAYERS) as u32;
let active = stack.rendered_count(reveal).clamp(1, MAX_LAYERS) as u32;
self.ensure_array(width, height, active)?;
let mut encoder = self
@@ -558,60 +747,143 @@ impl MaskPass {
label: Some("mask-encoder"),
});
for (slot, layer) in stack.active().enumerate().take(MAX_LAYERS) {
let field = match (&layer.source, labels) {
(MaskSource::Regions { .. }, None) => {
log::warn!(
"mask layer {} is a region mask with no segmentation loaded; skipping",
layer.id
);
continue;
}
(MaskSource::Regions { .. }, Some(f)) => f,
(_, _) => &self.placeholder,
};
for (slot, layer) in stack.rendered(reveal).enumerate().take(MAX_LAYERS) {
// **The path a mask with one part takes is the path every mask
// took before parts existed**: drawn straight into the layer's
// slice, cleared by the draw itself. Nothing about an unedited
// library's rendering changes, and the scratch texture is never
// allocated for it.
//
// An inverted base is the exception, because turning a part over
// is done where it is read back rather than where it is drawn —
// a brush deposits dabs and cannot know what the rest of the
// frame is. See `fs_combine`.
// TRACES: FR-DEV-19a
// The shown parts, not the parts: a hidden one is skipped here
// and nowhere else, and the first *shown* part is the one that
// opens the fold. Which can leave nothing — a revealed layer with
// every part hidden — and that clears the slice rather than
// leaving whatever the last rasterisation put there to be read
// back as this mask.
let shown: Vec<&dr_pipeline::mask::MaskPart> = layer.shown_parts().collect();
if shown.is_empty() {
self.clear_slice(&mut encoder, slot as u32);
continue;
}
let direct = shown.len() == 1 && !shown[0].invert;
if !direct {
self.ensure_scratch(width, height)?;
}
// A subject layer whose instance is missing is skipped for the
// same reason a region layer without a segmentation is: an absent
// mask that defaults to "everything" would apply the adjustment to
// the whole photograph, which is a much louder failure than none.
// Indexed by *slot*, not by the instance the layer names: the
// fields are built per layer, in this same order, because two
// layers over one subject can carry different morphology.
let subject = match &layer.source {
// Category alongside Subject: both are model coverage turned
// into a distance field, both are built per layer in this same
// order, and leaving a category out of here is precisely the
// failure the comment above warns about — it binds the 1x1
// placeholder, so the mask covers everything and the
// adjustment silently goes global.
MaskSource::Subject { .. } | MaskSource::Category { .. } => {
match subjects.filter(|s| slot < s.len()) {
Some(s) => (s, slot),
None => {
log::warn!("mask layer {} has no distance field; skipping", layer.id);
continue;
for (index, part) in shown.iter().copied().enumerate() {
let base = index == 0;
let field = match (&part.source, labels) {
(MaskSource::Regions { .. }, None) => {
log::warn!(
"mask layer {} is a region mask with no segmentation loaded; skipping",
layer.id
);
if base {
break;
}
continue;
}
(MaskSource::Regions { .. }, Some(f)) => f,
(_, _) => &self.placeholder,
};
// A subject part whose instance is missing is skipped for the
// same reason a region part without a segmentation is: an
// absent mask that defaults to "everything" would apply the
// adjustment to the whole photograph, which is a much louder
// failure than none.
//
// Indexed by *slot*, not by the instance the part names: the
// fields are built per layer, in this same order, because two
// layers over one subject can carry different morphology.
// Which is also why only a base part can have one — a model
// part joined to a mask has no field built for it yet, and it
// is skipped rather than drawn against a placeholder that
// would cover the frame.
let subject = match &part.source {
// Category alongside Subject: both are model coverage
// turned into a distance field, both are built per layer
// in this same order, and leaving a category out of here
// is precisely the failure the comment above warns about —
// it binds the 1x1 placeholder, so the mask covers
// everything and the adjustment silently goes global.
MaskSource::Subject { .. } | MaskSource::Category { .. } => {
match subjects.filter(|s| base && slot < s.len()) {
Some(s) => (s, slot),
None => {
log::warn!(
"part {} of mask layer {} has no distance field; skipping",
part.id,
layer.id
);
if base {
break;
}
continue;
}
}
}
}
_ => (&self.empty_subject, 0),
};
_ => (&self.empty_subject, 0),
};
let params = self.params(layer, field, width, height);
match &layer.source {
MaskSource::Brush { strokes } => {
self.draw_brush(&mut encoder, slot as u32, &params, strokes, width, height)
// TRACES: FR-DEV-10
// A range part with no photograph bound is skipped rather than
// drawn against the placeholder, on exactly the rule the two
// cases above follow: an absent mask that defaults to
// "everything" takes a local adjustment global, which is a far
// quieter failure than a part that visibly did not render.
let image = match (&part.source, source) {
(s, None) if s.is_range() => {
log::warn!(
"mask layer {} selects a range with no image loaded; skipping",
layer.id
);
if base {
break;
}
continue;
}
(_, image) => image,
};
let params = self.params(part, field, image, width, height);
let target = if direct {
self.slice_view(slot as u32)
} else {
self.scratch_view()
};
match &part.source {
MaskSource::Brush { strokes } => {
self.draw_brush(&mut encoder, &target, &params, strokes, width, height)
}
_ => {
let selected = self.selection_buffer(part, field);
self.draw(
&mut encoder,
&target,
&params,
field,
&selected,
subject,
image,
);
}
}
_ => {
let selected = self.selection_buffer(layer, field);
self.draw(
&mut encoder,
slot as u32,
&params,
field,
&selected,
subject,
);
if !direct {
// The first part joins a cleared slice, so it lands
// exactly as it was drawn whichever way it says it joins —
// there is nothing yet for a subtraction to take away
// from, and a mask that began by subtracting from nothing
// would render as empty however it was painted afterwards.
let join = if base { Join::Union } else { part.join };
self.combine(&mut encoder, slot as u32, join, base, &params);
}
}
}
@@ -632,11 +904,36 @@ impl MaskPass {
fn params(
&self,
layer: &dr_pipeline::mask::MaskLayer,
part: &dr_pipeline::mask::MaskPart,
field: &LabelField,
source: Option<&DemosaicedImage>,
width: u32,
height: u32,
) -> MaskParams {
// TRACES: FR-DEV-10
// How much of the photograph one mask texel covers. One when there is
// no image bound, which is a value nothing reads — the range modes are
// the only readers and they are skipped in that case.
let source_step = match source {
Some(image) => {
let (sw, sh) = image.size();
[
sw as f32 / width.max(1) as f32,
sh as f32 / height.max(1) as f32,
]
}
None => [1.0, 1.0],
};
// Row-major nine, widened to three `vec4`s. Identity where there is no
// image, so a range that somehow reached the shader without one would
// read camera values rather than nothing — the same defensive choice
// the demosaicer makes for an uncalibrated body.
let m = source.map_or([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0], |i| {
i.color_matrix()
});
let wb = source.map_or([1.0, 1.0, 1.0], |i| i.as_shot_wb());
let non_linear = source.is_some_and(|i| i.is_non_linear());
let base = MaskParams {
width,
height,
@@ -650,10 +947,18 @@ impl MaskPass {
axis: [1.0, 0.0],
softness: 0.0,
angle: 0.0,
_pad1: [0.0, 0.0],
source_step,
cam_to_srgb: [
[m[0], m[1], m[2], 0.0],
[m[3], m[4], m[5], 0.0],
[m[6], m[7], m[8], 0.0],
],
as_shot_wb: [wb[0], wb[1], wb[2], if non_linear { 1.0 } else { 0.0 }],
invert: u32::from(part.invert),
_pad: [0; 3],
};
match &layer.source {
match &part.source {
// `softness` carries the layer's feather. The model's coverage is
// already a soft sigmoid, so zero means "use the edge the model
// drew" rather than "hard edge" — the one place in this shader
@@ -673,12 +978,12 @@ impl MaskPass {
MaskParams {
mode: MODE_SUBJECT,
// `softness` is the feather half-width in pixels.
softness: (layer.feather * short).max(0.0),
softness: (part.feather * short).max(0.0),
// `angle` carries the morphology offset — reused rather
// than padded, since a subject layer has no ellipse to
// rotate.
angle: morph_offset(layer) * short,
falloff: falloff_code(layer.falloff),
angle: morph_offset(part) * short,
falloff: falloff_code(part.falloff),
..base
}
}
@@ -719,6 +1024,33 @@ impl MaskPass {
mode: MODE_BRUSH,
..base
},
// TRACES: FR-DEV-10
// A band, carried in the fields the gradients measure geometry
// in. Reused rather than given their own, and it is not a
// shortcut: `centre` and `axis` are two pairs of floats whose
// meaning has always been the mode's to decide, and a range that
// added four more would grow the uniform every other mask pays
// for. What matters is that nothing here is a *coordinate* — a
// range is not a function of position at all.
MaskSource::Luminance { lo, hi, softness } => MaskParams {
mode: MODE_LUMINANCE,
axis: [*lo, *hi],
softness: *softness,
..base
},
MaskSource::Colour {
hue,
hue_width,
chroma_lo,
chroma_hi,
softness,
} => MaskParams {
mode: MODE_COLOUR,
centre: [*hue, *hue_width],
axis: [*chroma_lo, *chroma_hi],
softness: *softness,
..base
},
}
}
@@ -737,7 +1069,7 @@ impl MaskPass {
fn draw_brush(
&self,
encoder: &mut wgpu::CommandEncoder,
slot: u32,
target: &wgpu::TextureView,
params: &MaskParams,
strokes: &[Stroke],
width: u32,
@@ -799,19 +1131,10 @@ impl MaskPass {
})
});
let array = self.array.as_ref().expect("array ensured by caller");
let view = array.texture.create_view(&wgpu::TextureViewDescriptor {
label: Some("mask-slice"),
dimension: Some(wgpu::TextureViewDimension::D2),
base_array_layer: slot,
array_layer_count: Some(1),
..Default::default()
});
let mut pass = encoder.begin_render_pass(&wgpu::RenderPassDescriptor {
label: Some("mask-brush-pass"),
color_attachments: &[Some(wgpu::RenderPassColorAttachment {
view: &view,
view: target,
depth_slice: None,
resolve_target: None,
ops: wgpu::Operations {
@@ -857,11 +1180,11 @@ impl MaskPass {
/// One byte-flag per region, or a single zero for a non-region layer.
fn selection_buffer(
&self,
layer: &dr_pipeline::mask::MaskLayer,
part: &dr_pipeline::mask::MaskPart,
field: &LabelField,
) -> wgpu::Buffer {
let mut flags = vec![0u32; field.region_count.max(1) as usize];
if let MaskSource::Regions { ids, .. } = &layer.source {
if let MaskSource::Regions { ids, .. } = &part.source {
for &id in ids {
if let Some(slot) = flags.get_mut(id as usize) {
*slot = 1;
@@ -882,11 +1205,12 @@ impl MaskPass {
fn draw(
&self,
encoder: &mut wgpu::CommandEncoder,
slot: u32,
target: &wgpu::TextureView,
params: &MaskParams,
field: &LabelField,
selected: &wgpu::Buffer,
subject: (&SubjectMasks, usize),
source: Option<&DemosaicedImage>,
) {
let params_buf = self
.ctx
@@ -924,25 +1248,23 @@ impl MaskPass {
}),
),
},
// TRACES: FR-DEV-10
wgpu::BindGroupEntry {
binding: 6,
resource: wgpu::BindingResource::TextureView(
source.map_or(&self.empty_image, |i| i.view()),
),
},
],
});
// The array slice is selected by the attachment rather than by a
// uniform the shader reads — one fewer value that can disagree with
// where the pass actually writes.
let array = self.array.as_ref().expect("array ensured by caller");
let view = array.texture.create_view(&wgpu::TextureViewDescriptor {
label: Some("mask-slice"),
dimension: Some(wgpu::TextureViewDimension::D2),
base_array_layer: slot,
array_layer_count: Some(1),
..Default::default()
});
let mut pass = encoder.begin_render_pass(&wgpu::RenderPassDescriptor {
label: Some("mask-pass"),
color_attachments: &[Some(wgpu::RenderPassColorAttachment {
view: &view,
view: target,
depth_slice: None,
resolve_target: None,
ops: wgpu::Operations {
@@ -963,6 +1285,166 @@ impl MaskPass {
pass.draw(0..3, 0..1);
}
/// A view of one layer's slice of the array.
fn slice_view(&self, slot: u32) -> wgpu::TextureView {
// The array slice is selected by the attachment rather than by a
// uniform the shader reads — one fewer value that can disagree with
// where the pass actually writes.
let array = self.array.as_ref().expect("array ensured by caller");
array.texture.create_view(&wgpu::TextureViewDescriptor {
label: Some("mask-slice"),
dimension: Some(wgpu::TextureViewDimension::D2),
base_array_layer: slot,
array_layer_count: Some(1),
..Default::default()
})
}
fn scratch_view(&self) -> wgpu::TextureView {
self.scratch
.as_ref()
.expect("scratch ensured by caller")
.texture
.create_view(&wgpu::TextureViewDescriptor {
label: Some("mask-part"),
..Default::default()
})
}
/// Blend the part sitting in [`Self::scratch`] into a layer's slice.
///
/// `first` clears the slice instead of loading it, which is both cheaper
/// on a tiler and the only thing that makes the fold start from nothing
/// covered rather than from whatever the last rasterisation left.
fn combine(
&self,
encoder: &mut wgpu::CommandEncoder,
slot: u32,
join: Join,
first: bool,
params: &MaskParams,
) {
let params_buf = self
.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("mask-combine-params"),
contents: bytemuck::bytes_of(params),
usage: wgpu::BufferUsages::UNIFORM,
});
let bind_group = self
.ctx
.device
.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("mask-combine-bind"),
layout: &self.combine_layout,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: params_buf.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 7,
resource: wgpu::BindingResource::TextureView(&self.scratch_view()),
},
],
});
let target = self.slice_view(slot);
let mut pass = encoder.begin_render_pass(&wgpu::RenderPassDescriptor {
label: Some("mask-combine-pass"),
color_attachments: &[Some(wgpu::RenderPassColorAttachment {
view: &target,
depth_slice: None,
resolve_target: None,
ops: wgpu::Operations {
load: if first {
wgpu::LoadOp::Clear(wgpu::Color::BLACK)
} else {
wgpu::LoadOp::Load
},
store: wgpu::StoreOp::Store,
},
})],
depth_stencil_attachment: None,
timestamp_writes: None,
occlusion_query_set: None,
multiview_mask: None,
});
pass.set_pipeline(match join {
Join::Union => &self.combine_union,
Join::Subtract => &self.combine_subtract,
});
pass.set_bind_group(0, &bind_group, &[]);
pass.draw(0..3, 0..1);
}
/// Leave a layer's slice covering nothing.
///
/// A pass that clears and draws nothing, for the one case where a layer
/// reaches the array with no part to draw: every part hidden while the
/// layer is being revealed. The slice has to be written, because the
/// shader reads it whatever this function did.
fn clear_slice(&self, encoder: &mut wgpu::CommandEncoder, slot: u32) {
let target = self.slice_view(slot);
encoder.begin_render_pass(&wgpu::RenderPassDescriptor {
label: Some("mask-clear-pass"),
color_attachments: &[Some(wgpu::RenderPassColorAttachment {
view: &target,
depth_slice: None,
resolve_target: None,
ops: wgpu::Operations {
load: wgpu::LoadOp::Clear(wgpu::Color::BLACK),
store: wgpu::StoreOp::Store,
},
})],
depth_stencil_attachment: None,
timestamp_writes: None,
occlusion_query_set: None,
multiview_mask: None,
});
}
/// The texture a part is drawn in before it is joined.
///
/// Allocated on the first mask that has more than one part and kept at the
/// rasterisation size, which is the same size the array is: a part and the
/// slice it joins are compared texel for texel, so there is nothing to
/// scale and nothing to sample between.
fn ensure_scratch(&mut self, width: u32, height: u32) -> Result<(), GpuError> {
if self
.scratch
.as_ref()
.is_some_and(|s| s.width == width && s.height == height)
{
return Ok(());
}
let texture = self.ctx.device.create_texture(&wgpu::TextureDescriptor {
label: Some("mask-scratch"),
size: wgpu::Extent3d {
width,
height,
depth_or_array_layers: 1,
},
mip_level_count: 1,
sample_count: 1,
dimension: wgpu::TextureDimension::D2,
format: MaskArray::FORMAT,
usage: wgpu::TextureUsages::RENDER_ATTACHMENT | wgpu::TextureUsages::TEXTURE_BINDING,
view_formats: &[],
});
self.scratch = Some(Scratch {
texture,
width,
height,
});
Ok(())
}
fn ensure_array(&mut self, width: u32, height: u32, layers: u32) -> Result<(), GpuError> {
if self
.array
@@ -1005,6 +1487,57 @@ impl MaskPass {
}
}
/// The texture one part is drawn into on its way into a layer's slice.
struct Scratch {
texture: wgpu::Texture,
width: u32,
height: u32,
}
/// `max(dst, src)` — the union of two parts.
///
/// Not source-over, which would build up: two parts that each half-cover a
/// pixel select it half, and adding them would make the overlap of two soft
/// edges harder than either of them, drawing a seam exactly where a
/// photographer joined two selections to avoid one.
const MAX_BLEND: wgpu::BlendComponent = wgpu::BlendComponent {
src_factor: wgpu::BlendFactor::One,
dst_factor: wgpu::BlendFactor::One,
operation: wgpu::BlendOperation::Max,
};
/// TRACES: FR-DEV-10
/// A single black texel, bound at the image slot for a mask that is not a
/// range.
///
/// Written rather than merely allocated. Undefined contents would be read by
/// nothing today, but a binding mistake in a range mask would then produce
/// whatever the driver left in memory — a mask that flickers between builds
/// and machines, which is the hardest shape of bug this pass could have.
fn empty_image(ctx: &GpuContext) -> wgpu::TextureView {
let texture = ctx.device.create_texture_with_data(
&ctx.queue,
&wgpu::TextureDescriptor {
label: Some("mask-empty-image"),
size: wgpu::Extent3d {
width: 1,
height: 1,
depth_or_array_layers: 1,
},
mip_level_count: 1,
sample_count: 1,
dimension: wgpu::TextureDimension::D2,
format: DemosaicedImage::FORMAT,
usage: wgpu::TextureUsages::TEXTURE_BINDING,
view_formats: &[],
},
wgpu::util::TextureDataOrder::LayerMajor,
// Four half-floats of zero. Rgba16Float, so eight bytes.
&[0u8; 8],
);
texture.create_view(&wgpu::TextureViewDescriptor::default())
}
/// The shorter edge of the space the mask is rasterised in.
///
/// Feather and morphology are stored as fractions of it, so the same edit is
@@ -1018,11 +1551,11 @@ fn field_short_edge(width: u32, height: u32) -> f32 {
/// Zero for closing and opening: those are folded into the field itself when
/// it is built, because their second half acts on a shape the original field
/// does not describe.
fn morph_offset(layer: &dr_pipeline::mask::MaskLayer) -> f32 {
fn morph_offset(part: &dr_pipeline::mask::MaskPart) -> f32 {
use dr_pipeline::mask::Morphology;
match layer.morphology {
Morphology::Dilate => layer.morph_radius,
Morphology::Erode => -layer.morph_radius,
match part.morphology {
Morphology::Dilate => part.morph_radius,
Morphology::Erode => -part.morph_radius,
Morphology::None | Morphology::Close | Morphology::Open => 0.0,
}
}
@@ -1131,7 +1664,7 @@ mod tests {
let mut pass = MaskPass::new(&ctx).expect("mask pass");
let array = pass
.render(&stack, Some(&field), None, w, h)
.render(&stack, Some(&field), None, None, w, h)
.expect("render");
assert_eq!(array.size(), (w, h));
assert_eq!(array.layers(), 1);
@@ -1153,7 +1686,7 @@ mod tests {
}));
let mut pass = MaskPass::new(&ctx).expect("mask pass");
assert!(pass.render(&stack, None, None, 8, 8).is_ok());
assert!(pass.render(&stack, None, None, None, 8, 8).is_ok());
}
#[test]
@@ -1176,7 +1709,9 @@ mod tests {
}));
let mut pass = MaskPass::new(&ctx).expect("mask pass");
let array = pass.render(&stack, None, None, 16, 16).expect("render");
let array = pass
.render(&stack, None, None, None, 16, 16)
.expect("render");
assert_eq!(array.layers(), 2, "one slice per active layer");
}
@@ -1186,11 +1721,11 @@ mod tests {
fn painted(gestures: &[Gesture]) -> MaskLayer {
let mut layer = lit(MaskSource::brush());
for (erase, radius, path) in gestures {
layer.begin_stroke(*erase, *radius, 0.5, 1.0);
layer.begin_stroke(0, *erase, *radius, 0.5, 1.0);
for &(x, y) in path {
layer.extend_stroke(x, y);
layer.extend_stroke(0, x, y);
}
layer.end_stroke();
layer.end_stroke(0);
}
layer
}
@@ -1208,7 +1743,9 @@ mod tests {
stack.push(painted(&[(false, 0.1, vec![(0.2, 0.2), (0.8, 0.8)])]));
let mut pass = MaskPass::new(&ctx).expect("mask pass");
let array = pass.render(&stack, None, None, 32, 32).expect("render");
let array = pass
.render(&stack, None, None, None, 32, 32)
.expect("render");
assert_eq!(array.layers(), 1);
}
@@ -1258,7 +1795,7 @@ mod tests {
};
let mut pass = MaskPass::new(&ctx).expect("mask pass");
let array = pass
.render(&MaskStack::new(), None, None, 8, 8)
.render(&MaskStack::new(), None, None, None, 8, 8)
.expect("render");
assert_eq!(
array.layers(),
@@ -1281,17 +1818,20 @@ mod tests {
}));
let mut pass = MaskPass::new(&ctx).expect("mask pass");
pass.render(&stack, None, None, 32, 32).expect("render");
pass.render(&stack, None, None, None, 32, 32)
.expect("render");
assert_eq!(pass.allocations(), 1);
pass.render(&stack, None, None, 32, 32).expect("render");
pass.render(&stack, None, None, None, 32, 32)
.expect("render");
assert_eq!(
pass.allocations(),
1,
"same size and layer count should not reallocate"
);
pass.render(&stack, None, None, 64, 64).expect("render");
pass.render(&stack, None, None, None, 64, 64)
.expect("render");
assert_eq!(pass.allocations(), 2, "a resize must reallocate");
}
}
+547
View File
@@ -0,0 +1,547 @@
//! TRACES: FR-MRG-10 | FR-MRG-11
//! The merge: source frames warped into an output surface, chunk by chunk.
//!
//! The per-pixel half of a panorama (FR-MRG-10), on the GPU: the warp of a
//! source tile into an output chunk, the weighted accumulation across
//! frames, and the resolve to sixteen-bit samples. The geometry it is
//! given — rotations, focal length, projection — is `dr-pano`'s, solved on
//! proxies before any full-resolution pixel exists (panorama.md §5), and
//! that is what makes this simple: every output pixel's source coordinates
//! are a closed-form function, so a chunk can be produced from the source
//! tiles that project into it and nothing else.
//!
//! # The loop
//!
//! ```text
//! for each band of rows of the output:
//! for each chunk across the band:
//! zero the accumulator
//! for each frame whose footprint meets the chunk:
//! the source rectangle the chunk needs, from the geometry
//! render it camera-linear through the pipeline (the tile)
//! warp the tile into the chunk, accumulate ← GPU
//! resolve the chunk to u16 ← GPU
//! copy it into the band
//! hand the band to the writer (one DNG strip)
//! ```
//!
//! No stage holds the composite (FR-MRG-11): the working set is one
//! chunk's accumulator, one tile, one band of u16 rows. The frame textures
//! are the caller's to provide and cache — `source` is asked for frame `k`
//! as it is needed, and a caller short of memory may demosaic on demand.
//!
//! # What is not here yet
//!
//! A feathered blend, not seams and a Laplacian pyramid: the weight is the
//! distance to the frame's edge, which hides exposure steps and small
//! misalignments and does not hide parallax. Gain is a scalar per frame
//! the caller supplies. Both are panorama.md §10's step 5, after the path
//! writes a file end to end.
use std::sync::Arc;
use dr_pano::bundle::Cameras;
use dr_pano::projection::{Bounds, Projection};
use wgpu::util::DeviceExt;
use crate::readback::await_mapping;
use crate::{AdjustPass, DemosaicedImage, GpuContext, GpuError};
/// One frame's part in the merge.
pub struct MergeFrame {
/// The frame's edit, for its lens corrections — the only part of an
/// edit the camera-space tap uses (FR-MRG-2).
pub graph: Arc<dr_pipeline::EditGraph>,
/// Multiplies the frame's samples, to bring its exposure to the
/// reference frame's. 1.0 for no correction.
pub gain: f32,
}
/// The output the merge produces.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct MergeOutput {
pub projection: Projection,
/// The projection's scale in output pixels: the cylinder's radius, the
/// plane's distance. The source focal length at full resolution gives
/// output pixels the size of source pixels at the centre.
pub scale: f64,
/// The rectangle of the projection to produce, centred coordinates.
pub bounds: Bounds,
/// Pixels over which a frame's weight ramps up from its edge.
pub feather: f32,
/// Chunk size: the unit of GPU work and of memory.
pub chunk: (u32, u32),
/// Multiplies a normalised sample (1.0 = white) to the sensor's scale.
pub sample_scale: f32,
}
impl MergeOutput {
pub fn width(&self) -> u32 {
self.bounds.width().ceil().max(1.0) as u32
}
pub fn height(&self) -> u32 {
self.bounds.height().ceil().max(1.0) as u32
}
}
/// A band of finished rows: `rows × width × 3` RGB `u16`, plus a coverage
/// mask (`true` where any frame reached the pixel).
pub struct Band<'a> {
pub first_row: u32,
pub rows: u32,
pub rgb: &'a [u16],
pub covered: &'a [bool],
}
#[repr(C)]
#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
struct WarpParams {
chunk_origin: [f32; 2],
chunk_size: [u32; 2],
projection: u32,
proj_scale: f32,
focal: f32,
gain: f32,
r0: [f32; 4],
r1: [f32; 4],
r2: [f32; 4],
frame_size: [f32; 2],
tile_origin: [f32; 2],
tile_size: [u32; 2],
feather: f32,
_pad: f32,
}
#[repr(C)]
#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
struct ResolveParams {
chunk_size: [u32; 2],
scale: f32,
_pad: f32,
}
/// The two pipelines and the chunk buffers.
pub struct MergePass {
ctx: GpuContext,
warp: wgpu::ComputePipeline,
warp_layout: wgpu::BindGroupLayout,
resolve: wgpu::ComputePipeline,
resolve_layout: wgpu::BindGroupLayout,
/// Accumulator and packed output for the current chunk size.
buffers: Option<(wgpu::Buffer, wgpu::Buffer, wgpu::Buffer, (u32, u32))>,
}
impl MergePass {
pub fn new(ctx: &GpuContext) -> Result<Self, GpuError> {
let module = ctx
.device
.create_shader_module(wgpu::ShaderModuleDescriptor {
label: Some("merge"),
source: wgpu::ShaderSource::Wgsl(include_str!("shaders/merge.wgsl").into()),
});
let uniform = |binding| wgpu::BindGroupLayoutEntry {
binding,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Buffer {
ty: wgpu::BufferBindingType::Uniform,
has_dynamic_offset: false,
min_binding_size: None,
},
count: None,
};
let storage = |binding, read_only| wgpu::BindGroupLayoutEntry {
binding,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Buffer {
ty: wgpu::BufferBindingType::Storage { read_only },
has_dynamic_offset: false,
min_binding_size: None,
},
count: None,
};
let warp_layout = ctx
.device
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
label: Some("merge-warp-bgl"),
entries: &[
uniform(0),
wgpu::BindGroupLayoutEntry {
binding: 1,
visibility: wgpu::ShaderStages::COMPUTE,
ty: wgpu::BindingType::Texture {
// Unfilterable: rgba32float, loaded by hand.
sample_type: wgpu::TextureSampleType::Float { filterable: false },
view_dimension: wgpu::TextureViewDimension::D2,
multisampled: false,
},
count: None,
},
storage(2, false),
],
});
let resolve_layout =
ctx.device
.create_bind_group_layout(&wgpu::BindGroupLayoutDescriptor {
label: Some("merge-resolve-bgl"),
entries: &[uniform(0), storage(1, true), storage(2, false)],
});
let pipeline = |name: &str, layout: &wgpu::BindGroupLayout| {
let pl = ctx
.device
.create_pipeline_layout(&wgpu::PipelineLayoutDescriptor {
label: Some(name),
bind_group_layouts: &[Some(layout)],
immediate_size: 0,
});
ctx.device
.create_compute_pipeline(&wgpu::ComputePipelineDescriptor {
label: Some(name),
layout: Some(&pl),
module: &module,
entry_point: Some(name),
compilation_options: Default::default(),
cache: None,
})
};
Ok(MergePass {
ctx: ctx.clone(),
warp: pipeline("warp", &warp_layout),
warp_layout,
resolve: pipeline("resolve", &resolve_layout),
resolve_layout,
buffers: None,
})
}
/// Allocate the chunk buffers for this size if the last ones differ.
fn ensure_buffers(&mut self, chunk: (u32, u32)) {
if self.buffers.as_ref().is_none_or(|b| b.3 != chunk) {
let n = u64::from(chunk.0) * u64::from(chunk.1);
let acc = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
label: Some("merge-acc"),
size: n * 16,
usage: wgpu::BufferUsages::STORAGE | wgpu::BufferUsages::COPY_DST,
mapped_at_creation: false,
});
let out = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
label: Some("merge-out"),
size: n * 8,
usage: wgpu::BufferUsages::STORAGE | wgpu::BufferUsages::COPY_SRC,
mapped_at_creation: false,
});
let read = self.ctx.device.create_buffer(&wgpu::BufferDescriptor {
label: Some("merge-read"),
size: n * 8,
usage: wgpu::BufferUsages::COPY_DST | wgpu::BufferUsages::MAP_READ,
mapped_at_creation: false,
});
self.buffers = Some((acc, out, read, chunk));
}
}
fn chunk_buffers(&self) -> (&wgpu::Buffer, &wgpu::Buffer, &wgpu::Buffer) {
let b = self.buffers.as_ref().expect("ensured by the caller");
(&b.0, &b.1, &b.2)
}
/// Produce the whole output, band by band, handing each finished band
/// to `sink`.
///
/// `cameras` are in **full-resolution source pixels** (`frame_size`),
/// with frame `k` corresponding to `frames[k]` and `source(k)`. `source`
/// supplies the demosaiced frame on demand and may cache as it sees fit.
#[allow(clippy::too_many_arguments)]
pub fn merge<S, F>(
&mut self,
adjust: &mut AdjustPass,
frames: &[MergeFrame],
cameras: &Cameras,
frame_size: (u32, u32),
output: &MergeOutput,
mut source: S,
mut sink: F,
mut cancelled: impl FnMut() -> bool,
) -> Result<(), GpuError>
where
S: FnMut(usize) -> Result<Arc<DemosaicedImage>, GpuError>,
F: FnMut(Band<'_>) -> Result<(), GpuError>,
{
let (out_w, out_h) = (output.width(), output.height());
let (cw, ch) = (output.chunk.0.max(8), output.chunk.1.max(8));
let (fw, fh) = (frame_size.0 as f64, frame_size.1 as f64);
let mut band_rgb = vec![0u16; (out_w * ch * 3) as usize];
let mut band_cov = vec![false; (out_w * ch) as usize];
let mut chunk_px: Vec<u32> = Vec::new();
let mut y = 0u32;
while y < out_h {
let rows = ch.min(out_h - y);
band_rgb.iter_mut().for_each(|v| *v = 0);
band_cov.iter_mut().for_each(|v| *v = false);
let mut x = 0u32;
while x < out_w {
if cancelled() {
return Err(GpuError::Readback("merge cancelled".into()));
}
let cols = cw.min(out_w - x);
let origin = (
output.bounds.min_u + f64::from(x),
output.bounds.min_v + f64::from(y),
);
self.zero_accumulator((cols, rows));
for (k, frame) in frames.iter().enumerate() {
let Some(rect) = source_rect(
output.projection,
output.scale,
cameras,
k,
origin,
(cols, rows),
(fw, fh),
) else {
continue;
};
let image = source(k)?;
// The tile: that rectangle of the frame, camera-linear,
// at 1:1.
let view = dr_pipeline::CropRect {
x: (rect.0 as f32) / fw as f32,
y: (rect.1 as f32) / fh as f32,
width: (rect.2 as f32) / fw as f32,
height: (rect.3 as f32) / fh as f32,
};
let shader = frame.graph.compose_camera_linear(view);
let tile = adjust.render_camera_linear(&image, &shader, rect.2, rect.3)?;
let r = cameras.rotations[k].transpose();
let params = WarpParams {
chunk_origin: [origin.0 as f32, origin.1 as f32],
chunk_size: [cols, rows],
projection: match output.projection {
Projection::Perspective => 0,
Projection::Cylindrical => 1,
Projection::Spherical => 2,
},
proj_scale: output.scale as f32,
focal: cameras.focal as f32,
gain: frame.gain,
r0: [r.0[0][0] as f32, r.0[0][1] as f32, r.0[0][2] as f32, 0.0],
r1: [r.0[1][0] as f32, r.0[1][1] as f32, r.0[1][2] as f32, 0.0],
r2: [r.0[2][0] as f32, r.0[2][1] as f32, r.0[2][2] as f32, 0.0],
frame_size: [fw as f32, fh as f32],
tile_origin: [rect.0 as f32, rect.1 as f32],
tile_size: [rect.2, rect.3],
feather: output.feather,
_pad: 0.0,
};
self.accumulate(&params, tile);
}
self.resolve_chunk((cols, rows), output.sample_scale, &mut chunk_px)?;
// Into the band.
for row in 0..rows as usize {
for col in 0..cols as usize {
let px = chunk_px[(row * cols as usize + col) * 2..][..2].to_vec();
let i = row * out_w as usize + (x as usize + col);
band_rgb[i * 3] = (px[0] & 0xFFFF) as u16;
band_rgb[i * 3 + 1] = (px[0] >> 16) as u16;
band_rgb[i * 3 + 2] = (px[1] & 0xFFFF) as u16;
band_cov[i] = (px[1] >> 16) != 0;
}
}
x += cols;
}
sink(Band {
first_row: y,
rows,
rgb: &band_rgb[..(out_w * rows * 3) as usize],
covered: &band_cov[..(out_w * rows) as usize],
})?;
y += rows;
}
Ok(())
}
fn zero_accumulator(&mut self, chunk: (u32, u32)) {
self.ensure_buffers(chunk);
let (acc, _, _) = self.chunk_buffers();
let n = u64::from(chunk.0) * u64::from(chunk.1) * 16;
let mut enc = self.ctx.device.create_command_encoder(&Default::default());
enc.clear_buffer(acc, 0, Some(n));
self.ctx.queue.submit(Some(enc.finish()));
}
fn accumulate(&mut self, params: &WarpParams, tile: &wgpu::Texture) {
let chunk = (params.chunk_size[0], params.chunk_size[1]);
let uniforms = self
.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("merge-warp-params"),
contents: bytemuck::bytes_of(params),
usage: wgpu::BufferUsages::UNIFORM,
});
let view = tile.create_view(&Default::default());
self.ensure_buffers(chunk);
let (acc, _, _) = self.chunk_buffers();
let bind = self
.ctx
.device
.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("merge-warp-bg"),
layout: &self.warp_layout,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: uniforms.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 1,
resource: wgpu::BindingResource::TextureView(&view),
},
wgpu::BindGroupEntry {
binding: 2,
resource: acc.as_entire_binding(),
},
],
});
let mut enc = self.ctx.device.create_command_encoder(&Default::default());
{
let mut pass = enc.begin_compute_pass(&Default::default());
pass.set_pipeline(&self.warp);
pass.set_bind_group(0, &bind, &[]);
pass.dispatch_workgroups(chunk.0.div_ceil(8), chunk.1.div_ceil(8), 1);
}
self.ctx.queue.submit(Some(enc.finish()));
}
fn resolve_chunk(
&mut self,
chunk: (u32, u32),
scale: f32,
out: &mut Vec<u32>,
) -> Result<(), GpuError> {
let params = ResolveParams {
chunk_size: [chunk.0, chunk.1],
scale,
_pad: 0.0,
};
let uniforms = self
.ctx
.device
.create_buffer_init(&wgpu::util::BufferInitDescriptor {
label: Some("merge-resolve-params"),
contents: bytemuck::bytes_of(&params),
usage: wgpu::BufferUsages::UNIFORM,
});
let n = u64::from(chunk.0) * u64::from(chunk.1);
self.ensure_buffers(chunk);
let (acc, packed, read) = self.chunk_buffers();
let bind = self
.ctx
.device
.create_bind_group(&wgpu::BindGroupDescriptor {
label: Some("merge-resolve-bg"),
layout: &self.resolve_layout,
entries: &[
wgpu::BindGroupEntry {
binding: 0,
resource: uniforms.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 1,
resource: acc.as_entire_binding(),
},
wgpu::BindGroupEntry {
binding: 2,
resource: packed.as_entire_binding(),
},
],
});
let mut enc = self.ctx.device.create_command_encoder(&Default::default());
{
let mut pass = enc.begin_compute_pass(&Default::default());
pass.set_pipeline(&self.resolve);
pass.set_bind_group(0, &bind, &[]);
pass.dispatch_workgroups(chunk.0.div_ceil(8), chunk.1.div_ceil(8), 1);
}
enc.copy_buffer_to_buffer(packed, 0, read, 0, n * 8);
self.ctx.queue.submit(Some(enc.finish()));
let slice = read.slice(..n * 8);
let (tx, rx) = std::sync::mpsc::channel();
slice.map_async(wgpu::MapMode::Read, move |r| {
let _ = tx.send(r);
});
await_mapping(&self.ctx, &rx)?;
{
let data = slice.get_mapped_range();
out.clear();
out.extend_from_slice(bytemuck::cast_slice::<u8, u32>(&data));
}
read.unmap();
Ok(())
}
}
/// The rectangle of frame `k` (x, y, w, h in source pixels) a chunk reads,
/// or `None` if the chunk sees nothing of the frame.
///
/// Walks the chunk's border, projects each point into the frame, and takes
/// the bounding box with a two-pixel margin for the bilinear fetch. The
/// border rather than the corners because under a cylinder or sphere the
/// extreme of a footprint is not at a corner.
fn source_rect(
projection: Projection,
scale: f64,
cameras: &Cameras,
k: usize,
origin: (f64, f64),
size: (u32, u32),
frame: (f64, f64),
) -> Option<(u32, u32, u32, u32)> {
let (w, h) = (f64::from(size.0), f64::from(size.1));
let steps = 16;
let mut min = (f64::MAX, f64::MAX);
let mut max = (f64::MIN, f64::MIN);
let mut any = false;
let mut visit = |u: f64, v: f64| {
let d = projection.to_direction(scale, u, v);
if let Some((x, y)) = cameras.project(k, d) {
let (x, y) = (x + frame.0 / 2.0, y + frame.1 / 2.0);
min = (min.0.min(x), min.1.min(y));
max = (max.0.max(x), max.1.max(y));
any = true;
}
};
for s in 0..=steps {
let t = f64::from(s) / f64::from(steps);
visit(origin.0 + w * t, origin.1);
visit(origin.0 + w * t, origin.1 + h);
visit(origin.0, origin.1 + h * t);
visit(origin.0 + w, origin.1 + h * t);
}
// The interior too, coarsely: a chunk can contain a frame entirely.
for i in 1..4 {
for j in 1..4 {
visit(
origin.0 + w * f64::from(i) / 4.0,
origin.1 + h * f64::from(j) / 4.0,
);
}
}
if !any {
return None;
}
let x0 = (min.0.floor() - 2.0).max(0.0);
let y0 = (min.1.floor() - 2.0).max(0.0);
let x1 = (max.0.ceil() + 2.0).min(frame.0);
let y1 = (max.1.ceil() + 2.0).min(frame.1);
if x1 <= x0 || y1 <= y0 {
return None;
}
Some((x0 as u32, y0 as u32, (x1 - x0) as u32, (y1 - y0) as u32))
}
+10 -4
View File
@@ -160,12 +160,18 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {
// Green is measured. Red and blue are interpolated from their own
// axis, with a correction from the green Laplacian.
//
// Malvar "G at R/B locations" kernels, transposed per axis:
// chroma along the row: (5c + 4(w1+e1) - (nw+ne+sw+se) - (n2+s2) + 0.5(w2+e2)) / 8
// Malvar "R at green in R row" kernel, and its transpose:
// chroma along the row: (5c + 4(w1+e1) - (nw+ne+sw+se) - (w2+e2) + 0.5(n2+s2)) / 8
//
// The -1 goes on the two greens *along* the chroma axis and the +0.5
// on the pair across it. Transposed, both kernels still sum to zero
// and reconstruct a flat patch exactly, but on an edge the correction
// at green sites is half strength and the false colour doubles: a
// blue/yellow zipper around every clipped highlight.
let along_row =
(5.0 * c + 4.0 * (w1 + e1) - diag1 - vert2 + 0.5 * horiz2) * 0.125;
(5.0 * c + 4.0 * (w1 + e1) - diag1 - horiz2 + 0.5 * vert2) * 0.125;
let along_col =
(5.0 * c + 4.0 * (n1 + s1) - diag1 - horiz2 + 0.5 * vert2) * 0.125;
(5.0 * c + 4.0 * (n1 + s1) - diag1 - vert2 + 0.5 * horiz2) * 0.125;
let red_horizontal = red_is_horizontal(gid.x, gid.y);
let r = select(along_col, along_row, red_horizontal);
+254 -2
View File
@@ -28,7 +28,8 @@ struct MaskParams {
label_width: u32,
label_height: u32,
// 0 = regions, 1 = linear, 2 = radial, 3 = subject, 4 = brush.
// 0 = regions, 1 = linear, 2 = radial, 3 = subject, 4 = brush,
// 5 = luminance range, 6 = colour range.
//
// A brush does not read this — it has its own entry points, because it is
// the one mask that is not a function of the whole frame — but it is set
@@ -50,10 +51,49 @@ struct MaskParams {
// Linear: (cos, sin) of the ramp direction. Radial: semi-axes.
axis: vec2<f32>,
// Linear: ramp width. Radial: edge falloff as a fraction of the radius.
// A range: the fade at each edge of its band, in the band's own units.
softness: f32,
// Radial only: rotation of the ellipse.
angle: f32,
_pad1: vec2<f32>,
// TRACES: FR-DEV-10
// How many source texels one mask texel spans, per axis.
//
// The mask array is rasterised at a proxy size and the photograph is not,
// so one texel here covers several there. A range mask is a function of
// pixel *values*, and point-sampling one source texel in four would make
// its edge follow the sensor's noise wherever the picture has fine
// texture — speckle that is then a mask, and therefore visible in the
// adjustment. Averaging the footprint is what makes the band land on the
// tone the area actually is.
source_step: vec2<f32>,
// Camera RGB → linear sRGB, one row each. Only a range reads these: it is
// the one mask that looks at the photograph, and a hue is the body's own
// primaries until this matrix has been applied — so the same stored arc
// would select a different set of colours on every make of sensor.
cam_to_srgb_0: vec4<f32>,
cam_to_srgb_1: vec4<f32>,
cam_to_srgb_2: vec4<f32>,
// rgb: as-shot white balance. w: non-zero when the source arrived
// gamma-encoded rather than linear.
as_shot_wb: vec4<f32>,
// Whether this part is turned over before it joins the mask.
//
// Read by `fs_combine` and by nothing else, deliberately. A brush deposits
// dabs onto an empty field and has no idea what the rest of the frame is,
// so a stroke shader cannot invert anything; doing it where the finished
// part is read back is the one place that works for every kind of source.
invert: u32,
// Three scalars rather than a `vec3<u32>`: a three-component vector is
// aligned to sixteen bytes in the uniform address space, so it would sit
// at offset 144 and make this struct 160 bytes against the Rust side's
// 144 — a mismatch wgpu reports as a binding too small for the shader,
// several layers away from the padding that caused it.
_pad0: u32,
_pad1: u32,
_pad2: u32,
}
@group(0) @binding(0) var<uniform> p: MaskParams;
@@ -74,6 +114,19 @@ struct MaskParams {
// shrinking and feathering free: each is arithmetic on this, so a slider moves
// a uniform instead of rebuilding a mask.
@group(0) @binding(3) var subject: texture_2d<f32>;
// TRACES: FR-DEV-10
// The photograph itself, as the demosaicer left it: camera RGB, unbalanced,
// with no edit applied. A 1x1 placeholder for every mask that is a shape,
// because the bindings are fixed and a second pipeline differing only in what
// it ignores would cost more than one texel.
//
// **The unedited image, and that is the design rather than an accident of
// pass order.** A band over the *edited* result would move as the edit was
// made: raising the highlights would change which pixels counted as
// highlights, so the slider would chase its own mask. Measuring what the
// camera recorded means the selection stays where the photographer put it
// while they work on it.
@group(0) @binding(6) var image: texture_2d<f32>;
// A full-screen triangle rather than a quad: three vertices instead of six,
// no shared edge for the rasteriser to crack along, and no vertex buffer.
@@ -221,6 +274,177 @@ fn subject_mask(uv: vec2<f32>) -> f32 {
}
}
// ---------------------------------------------------------------------------
// Range masks (FR-DEV-10)
// ---------------------------------------------------------------------------
//
// The masks that select by what a pixel *is* rather than by where it sits.
// Nothing below reads `frame_delta`, and that absence is the point: a range is
// not a function of position, so it cannot be stretched by an aspect ratio,
// cannot drift under a crop, and comes out the same at a proxy size and at an
// export because the only thing it depends on is the photograph's own values.
//
// The band arrives entirely in the fields the gradients use — `axis` is the
// pair of bounds, `centre` is a colour range's arc, `softness` is the fade —
// so a range costs nothing in the uniform beyond the image transform above.
// Display-encoded sRGB back to linear.
//
// A JPEG is uploaded with its bytes untouched, so its values are gamma-encoded
// where the demosaicer's are linear. The same undoing the generated adjust
// shader does, at the same point and for the same reason: a band over
// brightness is meaningless if two sources disagree about what a value means.
fn decode_srgb(c: vec3<f32>) -> vec3<f32> {
let lo = c / 12.92;
let hi = pow((max(c, vec3<f32>(0.04045)) + 0.055) / 1.055, vec3<f32>(2.4));
return select(hi, lo, c <= vec3<f32>(0.04045));
}
// One source texel, as linear sRGB.
//
// This is the prologue of the generated adjust shader, repeated: decode,
// balance, pull a clipped pixel back to neutral, then the camera matrix. It is
// repeated rather than shared because the composer emits WGSL for the *edit*
// and this pass is not one — but it must agree with it, since a range mask
// exists to select the values the layer's own adjustments will then see.
//
// The highlight desaturation is the part that looks skippable and is not. A
// fully clipped photosite arrives as (1,1,1), carrying no colour at all; the
// as-shot multipliers are far from neutral, so balancing it and passing it
// through the matrix produces a strong magenta. A colour range would then
// select every blown sky as if the photographer had asked for magenta.
fn source_texel(px: vec2<i32>) -> vec3<f32> {
var c = textureLoad(image, px, 0).rgb;
if (p.as_shot_wb.w > 0.5) {
c = decode_srgb(c);
}
let clipped = smoothstep(0.985, 1.0, max(c.r, max(c.g, c.b)));
c = c * p.as_shot_wb.rgb;
if (clipped > 0.0) {
c = mix(c, vec3<f32>(max(c.r, max(c.g, c.b))), clipped);
}
return vec3<f32>(
dot(p.cam_to_srgb_0.rgb, c),
dot(p.cam_to_srgb_1.rgb, c),
dot(p.cam_to_srgb_2.rgb, c),
);
}
// The most taps one mask texel averages, per axis.
//
// A cap rather than the true footprint. At a 1600 px proxy over a 24 MP frame
// the ratio is under four, so this is the whole footprint for every ordinary
// photograph; past it the taps stride across the footprint instead of
// covering it, which is a sample of the area rather than its mean. That is the
// right way to run out of budget here — the estimate gets noisier, it does not
// start measuring somewhere else.
const MAX_SOURCE_TAPS: i32 = 4;
// The photograph's value under one mask texel, in linear sRGB.
fn image_value(px: vec2<i32>) -> vec3<f32> {
let dims = vec2<i32>(textureDimensions(image));
let last = dims - vec2<i32>(1);
// The footprint's top-left corner in source texels. Not a centre plus a
// radius: the mask texel is a *box* over the source, and sampling
// symmetrically about its centre would weight the middle of every
// footprint twice at odd tap counts.
let origin = vec2<f32>(px) * p.source_step;
let taps = clamp(vec2<i32>(ceil(p.source_step)), vec2<i32>(1), vec2<i32>(MAX_SOURCE_TAPS));
// `stride`, not `step`: WGSL has a builtin of that name, and a local that
// shadows one is legal and unreadable in the same breath.
let stride = p.source_step / vec2<f32>(taps);
var total = vec3<f32>(0.0);
for (var y = 0; y < taps.y; y = y + 1) {
for (var x = 0; x < taps.x; x = x + 1) {
let at = origin + (vec2<f32>(f32(x), f32(y)) + vec2<f32>(0.5)) * stride;
total = total + source_texel(clamp(vec2<i32>(at), vec2<i32>(0), last));
}
}
return total / f32(taps.x * taps.y);
}
// A soft band: one inside, nothing outside, a smooth ramp across each edge.
//
// The `min` rather than a product of the two ramps. A band narrower than twice
// its softness has no plateau, and multiplying the rising and falling ramps
// would then peak well below one — so "select the highlights" would come out
// at sixty per cent and the photographer would compensate with opacity,
// against a mask that was quietly weaker than it said. `min` keeps the
// plateau where there is one and degrades to a single peak where there is not.
fn band(v: f32, lo: f32, hi: f32, soft: f32) -> f32 {
if (soft <= 0.0) {
return select(0.0, 1.0, v >= lo && v <= hi);
}
return min(smoothstep(lo - soft, lo, v), 1.0 - smoothstep(hi, hi + soft, v));
}
fn luminance_mask(px: vec2<i32>) -> f32 {
let y = dot(image_value(px), vec3<f32>(0.2126, 0.7152, 0.0722));
// Onto the perceptual position `tone_position` in `ops/_helpers.yaml`
// establishes, which is where the stored bounds are measured. Linear light
// puts middle grey at 0.18, so a band stated in it would spend four fifths
// of its travel inside the shadows.
let t = clamp(pow(max(y, 0.0), 1.0 / 3.0), 0.0, 1.0);
return band(t, p.axis.x, p.axis.y, p.softness);
}
// Hue in turns, 0 at red and increasing through yellow.
//
// The plain six-sector definition. Zero for a neutral, which is a value the
// caller must not act on — the chroma bound below is what keeps a colour range
// away from the greys where this number is rounding noise.
fn hue_of(c: vec3<f32>) -> f32 {
let hi = max(c.r, max(c.g, c.b));
let lo = min(c.r, min(c.g, c.b));
let d = hi - lo;
if (d <= 0.0) {
return 0.0;
}
var h = 0.0;
if (hi == c.r) {
h = (c.g - c.b) / d;
} else if (hi == c.g) {
h = (c.b - c.r) / d + 2.0;
} else {
h = (c.r - c.g) / d + 4.0;
}
return fract(h / 6.0);
}
fn colour_mask(px: vec2<i32>) -> f32 {
let c = max(image_value(px), vec3<f32>(0.0));
let hi = max(c.r, max(c.g, c.b));
let lo = min(c.r, min(c.g, c.b));
// The max-minus-min chroma `colour_saturation` uses, so the number the
// band is stated in is the one the rest of the pipeline means by
// 'colourfulness'.
var chroma = 0.0;
if (hi > 0.0) {
chroma = (hi - lo) / hi;
}
// Distance round the circle, so an arc centred near red reaches both ways
// past zero. Written as a wrap rather than as two comparisons because red
// is exactly where skin sits, and an arc that stopped at the seam would
// select half of it.
let d = abs(fract(hue_of(c) - p.centre.x + 0.5) - 0.5);
var arc = 0.0;
if (p.softness <= 0.0) {
arc = select(0.0, 1.0, d <= p.centre.y);
} else {
arc = 1.0 - smoothstep(p.centre.y, p.centre.y + p.softness, d);
}
// Both, not either: an arc alone selects a haze of noise everywhere the
// picture is nearly grey, because a hue rounded out of three almost-equal
// channels is still a hue.
return min(arc, band(chroma, p.axis.x, p.axis.y, p.softness));
}
@fragment
fn fs(@builtin(position) pos: vec4<f32>) -> @location(0) vec4<f32> {
let px = vec2<i32>(i32(pos.x), i32(pos.y));
@@ -235,6 +459,8 @@ fn fs(@builtin(position) pos: vec4<f32>) -> @location(0) vec4<f32> {
case 1u: { m = linear_mask(uv); }
case 2u: { m = radial_mask(uv); }
case 3u: { m = subject_mask(uv); }
case 5u: { m = luminance_mask(px); }
case 6u: { m = colour_mask(px); }
default: { m = 0.0; }
}
@@ -387,3 +613,29 @@ fn fs_brush(in: BrushVertex) -> @location(0) vec4<f32> {
return vec4<f32>(clamp(coverage * s.flow, 0.0, 1.0), 0.0, 0.0, 1.0);
}
// ---------------------------------------------------------------------------
// Joining one part to the mask so far
// ---------------------------------------------------------------------------
//
// A layer's mask is a fold over its parts, and the set operation is the *blend
// state* rather than arithmetic here: union is `max(dst, src)`, subtraction is
// `dst * (1 - src)`. Both are fixed-function, so joining a part costs one
// full-screen draw and no second texture beyond the one being read.
//
// # Why a part is drawn aside first, rather than straight onto the mask
//
// Because an erase stroke inside a part means "a hole in *this* part", not "a
// hole in the mask". Painted straight onto the accumulator it would take away
// whatever the parts before it had put there — so a correction that tidied its
// own edge would punch through the subject underneath, and the failure would
// look like the model's mask had holes in it.
@group(0) @binding(7) var part_mask: texture_2d<f32>;
@fragment
fn fs_combine(@builtin(position) pos: vec4<f32>) -> @location(0) vec4<f32> {
let v = textureLoad(part_mask, vec2<i32>(i32(pos.x), i32(pos.y)), 0).r;
let m = select(v, 1.0 - v, p.invert != 0u);
return vec4<f32>(clamp(m, 0.0, 1.0), 0.0, 0.0, 1.0);
}
+166
View File
@@ -0,0 +1,166 @@
// TRACES: FR-MRG-10 | FR-MRG-11
// The merge: one source tile warped into one output chunk, accumulated.
//
// Two entry points. `warp` runs once per (chunk, frame): for every chunk
// pixel it asks which direction that pixel looks along, turns the
// direction into the frame's camera, projects it to a source pixel, and
// if that pixel is inside the tile that was rendered for this chunk,
// samples it and adds it — weighted by its distance from the frame's edge
// — into the accumulator. `resolve` runs once per chunk after every frame
// has been added: divides the sums by the weights and packs the result as
// sixteen-bit samples at the sensor's scale (FR-MRG-3).
//
// The accumulator is a buffer and not a storage texture, because WebGPU
// allows a read-write storage texture only in the 32-bit single-channel
// formats, and this wants four channels. The tile is sampled by hand from
// four `textureLoad`s rather than through a sampler, because `rgba32float`
// is not filterable without an optional feature, and the tile is
// `rgba32float` on purpose (panorama.md §5.1).
//
// The projection maths is `dr_pano::projection` verbatim; the two must
// agree, and a golden test compares them.
struct Params {
// Where the chunk's pixel (0, 0) sits in centred output coordinates,
// and the chunk's size.
chunk_origin: vec2<f32>,
chunk_size: vec2<u32>,
// 0 perspective, 1 cylindrical, 2 spherical; and the projection's
// scale (the cylinder's radius, the sphere's, the plane's distance) in
// output pixels.
projection: u32,
proj_scale: f32,
// The frame's focal length in source pixels, and the gain the frame's
// exposure is corrected by.
focal: f32,
gain: f32,
// World → this frame's camera: the transpose of its rotation, one row
// per vec4 (padded).
r0: vec4<f32>,
r1: vec4<f32>,
r2: vec4<f32>,
// The full frame's size in source pixels (for the edge weight), the
// tile's origin within the frame, and the tile's size.
frame_size: vec2<f32>,
tile_origin: vec2<f32>,
tile_size: vec2<u32>,
// Pixels over which the weight ramps from the edge to full.
feather: f32,
_pad: f32,
};
@group(0) @binding(0) var<uniform> p: Params;
@group(0) @binding(1) var tile: texture_2d<f32>;
// rgb·w summed, then w: four floats per chunk pixel.
@group(0) @binding(2) var<storage, read_write> acc: array<vec4<f32>>;
fn to_direction(u: f32, v: f32) -> vec3<f32> {
let s = p.proj_scale;
if (p.projection == 0u) {
return normalize(vec3<f32>(u, v, s));
}
if (p.projection == 1u) {
let theta = u / s;
return normalize(vec3<f32>(sin(theta), v / s, cos(theta)));
}
let theta = u / s;
let phi = v / s;
return vec3<f32>(sin(theta) * cos(phi), sin(phi), cos(theta) * cos(phi));
}
fn load(x: i32, y: i32) -> vec4<f32> {
return textureLoad(tile, vec2<i32>(x, y), 0);
}
@compute @workgroup_size(8, 8, 1)
fn warp(@builtin(global_invocation_id) gid: vec3<u32>) {
if (gid.x >= p.chunk_size.x || gid.y >= p.chunk_size.y) {
return;
}
let u = p.chunk_origin.x + f32(gid.x) + 0.5;
let v = p.chunk_origin.y + f32(gid.y) + 0.5;
let d = to_direction(u, v);
let c = vec3<f32>(dot(p.r0.xyz, d), dot(p.r1.xyz, d), dot(p.r2.xyz, d));
if (c.z <= 1e-6) {
return;
}
// Source pixel, in the full frame, with the principal point at its
// centre. `- 0.5` puts pixel centres on integer coordinates for the
// bilinear fetch below.
let sx = p.focal * c.x / c.z + p.frame_size.x * 0.5 - 0.5;
let sy = p.focal * c.y / c.z + p.frame_size.y * 0.5 - 0.5;
// Weight: distance to the nearest frame edge, in pixels, over the
// feather. Zero outside the frame.
let edge = min(min(sx, p.frame_size.x - 1.0 - sx), min(sy, p.frame_size.y - 1.0 - sy));
if (edge <= 0.0) {
return;
}
let w = clamp(edge / max(p.feather, 1.0), 0.0, 1.0);
// Into the tile.
let tx = sx - p.tile_origin.x;
let ty = sy - p.tile_origin.y;
let tw = f32(p.tile_size.x);
let th = f32(p.tile_size.y);
if (tx < 0.0 || ty < 0.0 || tx > tw - 1.0 || ty > th - 1.0) {
return;
}
let x0 = i32(floor(tx));
let y0 = i32(floor(ty));
let x1 = min(x0 + 1, i32(p.tile_size.x) - 1);
let y1 = min(y0 + 1, i32(p.tile_size.y) - 1);
let fx = tx - f32(x0);
let fy = ty - f32(y0);
// The four texels, with their alpha: the tap writes alpha 0 where the
// lens correction found no source pixel, and a sample that touches one
// of those is a partial pixel — down-weighted by exactly how much of
// it is missing, and dropped when all of it is.
let s00 = load(x0, y0);
let s10 = load(x1, y0);
let s01 = load(x0, y1);
let s11 = load(x1, y1);
let top = mix(s00, s10, fx);
let bot = mix(s01, s11, fx);
let s = mix(top, bot, fy);
if (s.a <= 0.001) {
return;
}
// Colour is the alpha-weighted mean of the texels that exist.
let rgb = s.rgb / s.a * p.gain;
let wa = w * s.a;
let i = gid.y * p.chunk_size.x + gid.x;
acc[i] = acc[i] + vec4<f32>(rgb * wa, wa);
}
// Resolve: the accumulated chunk to sixteen-bit samples.
struct ResolveParams {
chunk_size: vec2<u32>,
// Multiplies a normalised value (1.0 = the sensor's white) back to the
// sensor's scale: the source's white minus its black (FR-MRG-3).
scale: f32,
_pad: f32,
};
@group(0) @binding(0) var<uniform> rp: ResolveParams;
@group(0) @binding(1) var<storage, read> racc: array<vec4<f32>>;
// Two u32 per pixel: (r | g << 16), (b | coverage << 16). Coverage is
// 65535 where any frame reached the pixel and 0 where none did, so the
// CPU can tell an empty pixel from a black one.
@group(0) @binding(2) var<storage, read_write> out: array<vec2<u32>>;
@compute @workgroup_size(8, 8, 1)
fn resolve(@builtin(global_invocation_id) gid: vec3<u32>) {
if (gid.x >= rp.chunk_size.x || gid.y >= rp.chunk_size.y) {
return;
}
let i = gid.y * rp.chunk_size.x + gid.x;
let a = racc[i];
if (a.w <= 0.0) {
out[i] = vec2<u32>(0u, 0u);
return;
}
let rgb = clamp(a.rgb / a.w * rp.scale, vec3<f32>(0.0), vec3<f32>(65535.0));
let r = u32(round(rgb.r));
let g = u32(round(rgb.g));
let b = u32(round(rgb.b));
out[i] = vec2<u32>(r | (g << 16u), b | (65535u << 16u));
}
+4
View File
@@ -40,6 +40,10 @@ fn flat_raw(level: u16, curve: BaseCurve) -> RawImage {
wb_coeffs: [1.0, 1.0, 1.0, 1.0],
color_matrix: Some([1.0, 0.0, 0.0, 0.0, 1.0, 0.0, 0.0, 0.0, 1.0]),
base_curve: curve,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
+4
View File
@@ -46,6 +46,10 @@ fn flat_raw(level: u16) -> RawImage {
// leaving a curve here would test the suppression rather than the
// film. `dr-pipeline` asserts the suppression on the generated source.
base_curve: BaseCurve::IDENTITY,
samples_per_pixel: 1,
profile: None,
make: String::new(),
model: String::new(),
crop: CropRect {
x: 0,
y: 0,
+426 -9
View File
@@ -12,8 +12,8 @@
use dr_gpu::{AdjustPass, DemosaicedImage, GpuContext, LabelField, MaskPass};
use dr_pipeline::descriptor::ParamId;
use dr_pipeline::mask::{MaskLayer, MaskSource, MaskStack};
use dr_pipeline::operation::compose_full;
use dr_pipeline::mask::{Join, MaskLayer, MaskPart, MaskSource, MaskStack, Reveal, RevealStyle};
use dr_pipeline::operation::{compose_full, compose_full_revealing};
use dr_pipeline::spot::SpotSet;
use dr_pipeline::{ops, EditGraph, Framing};
use dr_types::ColourSpace;
@@ -72,10 +72,13 @@ fn render_at(
ColourSpace::Srgb,
stack,
&SpotSet::new(),
&[],
);
let mut masks = MaskPass::new(ctx).expect("mask pass");
let array = masks.render(stack, field, None, w, h).expect("rasterise");
let array = masks
.render(stack, field, None, None, w, h)
.expect("rasterise");
let mut adjust = AdjustPass::new(ctx);
adjust
@@ -283,11 +286,11 @@ type Gesture = (bool, f32, f32, Vec<(f32, f32)>);
fn painted(gestures: &[Gesture]) -> MaskLayer {
let mut layer = brighten(MaskSource::brush());
for (erase, radius, flow, path) in gestures {
layer.begin_stroke(*erase, *radius, 0.9, *flow);
layer.begin_stroke(0, *erase, *radius, 0.9, *flow);
for &(x, y) in path {
layer.extend_stroke(x, y);
layer.extend_stroke(0, x, y);
}
layer.end_stroke();
layer.end_stroke(0);
}
layer
}
@@ -502,9 +505,9 @@ fn hardness_decides_how_quickly_the_edge_falls_away() {
let edge = |hardness: f32| {
let mut layer = brighten(MaskSource::brush());
layer.begin_stroke(false, 0.4, hardness, 1.0);
layer.extend_stroke(0.5, 0.5);
layer.end_stroke();
layer.begin_stroke(0, false, 0.4, hardness, 1.0);
layer.extend_stroke(0, 0.5, 0.5);
layer.end_stroke(0);
let pixels = render(&ctx, &stack_of(layer), None);
// How many pixels along the centre row are neither fully painted nor
@@ -605,3 +608,417 @@ fn an_empty_stack_renders_exactly_as_the_unmasked_path() {
let masked = render(&ctx, &MaskStack::new(), None);
assert_eq!(plain, masked, "an empty mask stack must be a no-op");
}
// ---------------------------------------------------------------------------
// Parts — a mask built from more than one selection
// ---------------------------------------------------------------------------
/// Everything, so a part joined to it has something to change.
fn whole_frame() -> MaskSource {
MaskSource::Regions {
signature: 1,
level: 2,
ids: vec![0, 1],
}
}
/// Paint one gesture into a part of a layer, at a radius large enough that a
/// 32-pixel frame can tell where it landed.
fn paint(layer: &mut MaskLayer, part: usize, erase: bool, path: &[(f32, f32)]) {
assert!(
layer.begin_stroke(part, erase, 0.2, 1.0, 1.0),
"the layer had no room for a stroke"
);
for &(x, y) in path {
layer.extend_stroke(part, x, y);
}
layer.end_stroke(part);
}
#[test]
fn a_union_part_adds_what_the_base_did_not_cover() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let mut layer = brighten(MaskSource::Regions {
signature: 1,
level: 2,
ids: vec![0],
});
assert!(layer.push_part(MaskPart::painted("p2", Join::Union)));
paint(&mut layer, 1, false, &[(0.85, 0.5)]);
let mut stack = MaskStack::new();
stack.push(layer);
let pixels = render(&ctx, &stack, Some(&split_field(&ctx)));
assert!(
luma_at(&pixels, 4, 16) > 200,
"the base region is still masked"
);
assert!(
luma_at(&pixels, 27, 16) > 200,
"and the painted part joined the other half in"
);
assert_eq!(
luma_at(&pixels, 20, 2),
128,
"while what neither covers is untouched"
);
}
#[test]
fn a_subtract_part_takes_a_bite_out_of_the_mask() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let mut layer = brighten(whole_frame());
assert!(layer.push_part(MaskPart::painted("p2", Join::Subtract)));
paint(&mut layer, 1, false, &[(0.5, 0.5)]);
let mut stack = MaskStack::new();
stack.push(layer);
let pixels = render(&ctx, &stack, Some(&split_field(&ctx)));
assert_eq!(
luma_at(&pixels, 16, 16),
128,
"the middle was taken back out of the mask"
);
assert!(
luma_at(&pixels, 1, 1) > 200,
"and the corner the stroke never reached is still in it"
);
}
/// **The test the scratch texture exists for.** An erase stroke means a hole in
/// the part it was painted into, not a hole in the mask: drawn straight onto
/// the accumulator it would take away whatever the parts before it had put
/// there, so tidying the edge of a correction would punch through the subject
/// underneath — and the failure would read as the model's mask having holes.
#[test]
fn an_erase_stroke_holes_its_own_part_and_not_the_mask() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let mut layer = brighten(whole_frame());
assert!(layer.push_part(MaskPart::painted("p2", Join::Union)));
paint(&mut layer, 1, false, &[(0.5, 0.5)]);
paint(&mut layer, 1, true, &[(0.5, 0.5)]);
let mut stack = MaskStack::new();
stack.push(layer);
let pixels = render(&ctx, &stack, Some(&split_field(&ctx)));
assert!(
luma_at(&pixels, 16, 16) > 200,
"the base still covers where the correction erased itself"
);
}
#[test]
fn an_inverted_part_joins_everything_it_did_not_paint() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
// A base that covers nothing, so what shows is the part alone.
let mut layer = brighten(MaskSource::brush());
paint(&mut layer, 0, false, &[(0.1, 0.1)]);
assert!(layer.push_part(MaskPart::painted("p2", Join::Union)));
layer.parts_mut()[1].invert = true;
paint(&mut layer, 1, false, &[(0.5, 0.5)]);
let mut stack = MaskStack::new();
stack.push(layer);
let pixels = render(&ctx, &stack, Some(&split_field(&ctx)));
assert_eq!(
luma_at(&pixels, 16, 16),
128,
"where the inverted part was painted is outside the mask"
);
assert!(
luma_at(&pixels, 30, 30) > 200,
"and everywhere it was not is inside it"
);
}
/// Parts fold in order, so the same two selections joined the other way round
/// are a different mask. A subtraction that lands before the part it was meant
/// to cut into would take away nothing at all.
#[test]
fn the_order_parts_are_joined_in_is_the_mask() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let render_pair = |cut_first: bool| {
let mut layer = brighten(MaskSource::brush());
if cut_first {
layer.push_part(MaskPart::painted("p2", Join::Subtract));
layer.push_part(MaskPart::painted("p3", Join::Union));
paint(&mut layer, 1, false, &[(0.5, 0.5)]);
paint(&mut layer, 2, false, &[(0.5, 0.5)]);
} else {
layer.push_part(MaskPart::painted("p2", Join::Union));
layer.push_part(MaskPart::painted("p3", Join::Subtract));
paint(&mut layer, 1, false, &[(0.5, 0.5)]);
paint(&mut layer, 2, false, &[(0.5, 0.5)]);
}
let mut stack = MaskStack::new();
stack.push(layer);
render(&ctx, &stack, Some(&split_field(&ctx)))
};
assert!(
luma_at(&render_pair(true), 16, 16) > 200,
"adding after a subtraction leaves the addition standing"
);
assert_eq!(
luma_at(&render_pair(false), 16, 16),
128,
"and subtracting after an addition takes it away again"
);
}
// --- seeing the mask (FR-DEV-19c) ------------------------------------------
/// A radial that covers the middle of the frame and nothing near the corners.
fn middle() -> MaskSource {
MaskSource::Radial {
centre: (0.5, 0.5),
radii: (0.3, 0.3),
angle: 0.0,
feather: 0.05,
}
}
/// [`render`], with one layer's mask drawn over the result.
fn render_revealing(ctx: &GpuContext, stack: &MaskStack, reveal: &Reveal) -> Vec<u8> {
let source = grey(ctx);
let shader = compose_full_revealing(
&ops::chain(),
&Framing::new(),
ColourSpace::Srgb,
stack,
&SpotSet::new(),
&[],
Some(reveal),
);
let mut masks = MaskPass::new(ctx).expect("mask pass");
let array = masks
.render_revealing(stack, None, None, None, SIZE, SIZE, Some(reveal))
.expect("rasterise");
let mut adjust = AdjustPass::new(ctx);
adjust
.render_masked(&source, &shader, SIZE, SIZE, Some(array))
.expect("render");
adjust.export_pixels().expect("readback").0
}
/// TRACES: FR-DEV-19c
/// The state every mask is in for its first few seconds: chosen, and not yet
/// used for anything.
///
/// Such a layer changes no pixel, so it is not active, so it occupied no mask
/// slot and was never rasterised — and the reveal drew nothing. That is the
/// whole of "I clicked the category and nothing happened": there was a mask,
/// and no way to see that there was.
#[test]
fn a_selection_with_no_adjustment_can_still_be_seen() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let mut stack = MaskStack::new();
stack.push(MaskLayer::new("m1", middle()));
assert!(
stack.is_neutral(),
"the fixture must be a selection with nothing done to it"
);
let pixels = render_revealing(&ctx, &stack, &Reveal::one("m1", RevealStyle::Alpha));
assert!(
luma_at(&pixels, SIZE / 2, SIZE / 2) > 200,
"the middle is inside the mask and should read white"
);
assert!(
luma_at(&pixels, 1, 1) < 40,
"the corner is outside it and should read black"
);
}
/// And with nobody looking, the same stack changes nothing at all.
///
/// The other half of the property above: a layer renders *because* it is being
/// revealed, so it must stop when the reveal does — otherwise a selection with
/// no adjustment would leave a slice in the array for ever.
#[test]
fn a_mask_nobody_is_looking_at_draws_nothing() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let mut stack = MaskStack::new();
stack.push(MaskLayer::new("m1", middle()));
let pixels = render(&ctx, &stack, None);
assert_eq!(
luma_at(&pixels, SIZE / 2, SIZE / 2),
128,
"flat grey, exactly as it went in"
);
}
/// TRACES: FR-DEV-19c
/// A tint has to leave the photograph visible, or it cannot be judged against
/// it — which is the one thing an overlay exists for.
#[test]
fn a_tint_colours_the_mask_and_leaves_the_rest_alone() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let mut stack = MaskStack::new();
stack.push(MaskLayer::new("m1", middle()));
let pixels = render_revealing(&ctx, &stack, &Reveal::one("m1", RevealStyle::Tint));
let at = |x: u32, y: u32| {
let i = ((y * SIZE + x) * 4) as usize;
(pixels[i], pixels[i + 1], pixels[i + 2])
};
let (r, g, _) = at(SIZE / 2, SIZE / 2);
assert!(r > g + 40, "the mask should read red, got r={r} g={g}");
assert!(
g > 20,
"and not opaque — the photograph under it is what the tint is judged \
against, got g={g}"
);
let (r, g, b) = at(1, 1);
assert!(
(120..=136).contains(&r) && r == g && g == b,
"outside the mask the photograph is untouched, got ({r}, {g}, {b})"
);
}
/// TRACES: FR-DEV-19c
/// An outline draws where the mask stops and nowhere else — which is the
/// point of it, since the other two styles cover the detail the boundary has
/// to be judged against.
#[test]
fn an_outline_draws_the_boundary_and_not_the_interior() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let mut stack = MaskStack::new();
stack.push(MaskLayer::new("m1", middle()));
let pixels = render_revealing(&ctx, &stack, &Reveal::one("m1", RevealStyle::Edge));
// Where the line landed, along the row through the centre. Searched
// rather than sampled at one place: the radial's edge crosses this row
// about 9.6 pixels out from the middle on a 32px frame, and asserting a
// particular pixel would be asserting the rounding.
let (at, brightest) = (SIZE / 2..SIZE)
.map(|x| (x, luma_at(&pixels, x, SIZE / 2)))
.max_by_key(|&(_, v)| v)
.expect("the row is not empty");
assert!(
brightest > 160,
"there should be a line somewhere on this row, brightest was {brightest}"
);
assert!(
(SIZE / 2 + 7..=SIZE / 2 + 12).contains(&at),
"and it should be on the mask's boundary, not somewhere else: x={at}"
);
assert_eq!(
luma_at(&pixels, SIZE / 2, SIZE / 2),
128,
"the picture inside the mask is untouched"
);
assert_eq!(
luma_at(&pixels, 1, 1),
128,
"and so is the picture outside it"
);
}
/// TRACES: FR-DEV-19c
/// Two masks shown at once come out in two colours, each where its own mask
/// is — which is what makes "where do these meet" a question the screen can
/// answer.
#[test]
fn two_shown_masks_are_drawn_each_in_its_own_colour() {
use dr_pipeline::mask::RevealedLayer;
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
// A left half and a right half, as two brush layers with one fat dab each.
let half = |id: &str, x: f32| {
let mut layer = MaskLayer::new(id, MaskSource::brush());
paint(&mut layer, 0, false, &[(x, 0.5)]);
layer
};
let mut stack = MaskStack::new();
stack.push(half("left", 0.2));
stack.push(half("right", 0.8));
let reveal = Reveal {
layers: vec![
RevealedLayer {
layer: "left".into(),
colour: [1.0, 0.0, 0.0],
},
RevealedLayer {
layer: "right".into(),
colour: [0.0, 0.0, 1.0],
},
],
style: RevealStyle::Alpha,
};
let pixels = render_revealing(&ctx, &stack, &reveal);
let at = |x: u32| {
let i = ((SIZE / 2 * SIZE + x) * 4) as usize;
(pixels[i], pixels[i + 1], pixels[i + 2])
};
let (r, _, b) = at(SIZE / 5);
assert!(
r > 200 && b < 40,
"the left mask reads red, got r={r} b={b}"
);
let (r, _, b) = at(SIZE * 4 / 5);
assert!(
b > 200 && r < 40,
"the right mask reads blue, got r={r} b={b}"
);
let (r, g, b) = at(SIZE / 2);
assert!(
r < 40 && g < 40 && b < 40,
"between them, alpha shows black: ({r}, {g}, {b})"
);
}
+8 -3
View File
@@ -51,7 +51,7 @@ fn brightening_stack() -> MaskStack {
},
);
layer.set_param("exposure", ParamId("exposure"), 2.0);
layer.feather = 0.0;
layer.base_mut().feather = 0.0;
stack.push(layer);
stack
}
@@ -81,7 +81,7 @@ fn render_at(ctx: &GpuContext, stack: &MaskStack, out: u32) -> Vec<u8> {
let mut masks = MaskPass::new(ctx).expect("mask pass");
let array = masks
.render(stack, None, Some(&subjects), PROXY, PROXY)
.render(stack, None, Some(&subjects), None, PROXY, PROXY)
.expect("rasterise");
let shader = compose_full(
@@ -90,6 +90,7 @@ fn render_at(ctx: &GpuContext, stack: &MaskStack, out: u32) -> Vec<u8> {
ColourSpace::Srgb,
stack,
&SpotSet::new(),
&[],
);
let mut adjust = AdjustPass::new(ctx);
adjust
@@ -108,6 +109,7 @@ fn render_unmasked(ctx: &GpuContext, stack: &MaskStack, out: u32) -> Vec<u8> {
ColourSpace::Srgb,
stack,
&SpotSet::new(),
&[],
);
let mut adjust = AdjustPass::new(ctx);
adjust.render(&source, &shader, out, out).expect("render");
@@ -222,7 +224,9 @@ fn render_gradient(ctx: &GpuContext, stack: &MaskStack, w: u32, h: u32) -> Vec<u
let source = DemosaicedImage::from_rgba8(ctx, &data, w, h).expect("upload");
let mut masks = MaskPass::new(ctx).expect("mask pass");
let array = masks.render(stack, None, None, w, h).expect("rasterise");
let array = masks
.render(stack, None, None, None, w, h)
.expect("rasterise");
let shader = compose_full(
&ops::chain(),
@@ -230,6 +234,7 @@ fn render_gradient(ctx: &GpuContext, stack: &MaskStack, w: u32, h: u32) -> Vec<u
ColourSpace::Srgb,
stack,
&SpotSet::new(),
&[],
);
let mut adjust = AdjustPass::new(ctx);
adjust
+200
View File
@@ -0,0 +1,200 @@
// TRACES: FR-DEV-10
//! A range mask must select by the photograph's values, and by nothing else.
//!
//! Two properties, and both are silent when they break. A range mask that
//! reads the wrong pixels still produces a plausible-looking selection, and one
//! that depends on the size it was rasterised at looks right in the develop
//! view and wrong only in the export — the one place nobody is watching.
//!
//! Everything here goes through the real pass. There is no CPU rasterisation of
//! a mask to test against and there must not be (ARCH §5.4), so the mask is
//! observed the only way it exists: through the adjustment it weights.
use dr_gpu::{AdjustPass, DemosaicedImage, GpuContext, MaskPass};
use dr_pipeline::descriptor::ParamId;
use dr_pipeline::mask::{MaskLayer, MaskSource, MaskStack};
use dr_pipeline::operation::compose_full;
use dr_pipeline::spot::SpotSet;
use dr_pipeline::{ops, Framing};
use dr_types::ColourSpace;
/// The source and the output are the same size, so a difference between two
/// runs can only have come from the mask.
const SIZE: u32 = 64;
fn ctx() -> Option<GpuContext> {
pollster::block_on(GpuContext::new_headless()).ok()
}
/// A frame whose left half is one colour and right half another.
///
/// Deliberately not a ramp. A range mask's whole job is to divide the picture
/// by value, so a source that is already divided by value makes the assertion
/// "the mask found the half it was aimed at" rather than "the mask is roughly
/// where it should be".
fn split(ctx: &GpuContext, left: [u8; 3], right: [u8; 3]) -> DemosaicedImage {
let data: Vec<u8> = (0..SIZE * SIZE)
.flat_map(|i| {
let c = if i % SIZE < SIZE / 2 { left } else { right };
[c[0], c[1], c[2], 255]
})
.collect();
DemosaicedImage::from_rgba8(ctx, &data, SIZE, SIZE).expect("upload")
}
fn brightened(source: MaskSource) -> MaskStack {
let mut stack = MaskStack::new();
let mut layer = MaskLayer::new("m1", source);
// Two stops, so "selected" and "not selected" are not a judgement call.
layer.set_param("exposure", ParamId("exposure"), 2.0);
stack.push(layer);
stack
}
/// Render `stack` over `source`, with the mask rasterised at `raster`.
///
/// `raster` is a parameter because it is the thing that must not matter: the
/// develop view and an export ask for the same mask at different sizes, and a
/// range that answered differently at each would be a mask that changes when
/// the photograph is exported.
fn render(ctx: &GpuContext, source: &DemosaicedImage, stack: &MaskStack, raster: u32) -> Vec<u8> {
let mut masks = MaskPass::new(ctx).expect("mask pass");
let array = masks
.render(stack, None, None, Some(source), raster, raster)
.expect("rasterise");
let shader = compose_full(
&ops::chain(),
&Framing::new(),
ColourSpace::Srgb,
stack,
&SpotSet::new(),
&[],
);
let mut adjust = AdjustPass::new(ctx);
adjust
.render_masked(source, &shader, SIZE, SIZE, Some(array))
.expect("render");
adjust.export_pixels().expect("readback").0
}
fn red_at(pixels: &[u8], x: u32, y: u32) -> u8 {
pixels[((y * SIZE + x) * 4) as usize]
}
/// The centre of each half, away from the seam the mask's own softness
/// straddles.
fn halves(pixels: &[u8]) -> (u8, u8) {
(
red_at(pixels, SIZE / 4, SIZE / 2),
red_at(pixels, SIZE * 3 / 4, SIZE / 2),
)
}
/// TRACES: FR-DEV-10
#[test]
fn a_tone_band_brightens_only_the_half_inside_it() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
// 200 lands near 0.83 on the perceptual scale and 30 near 0.24, so a band
// over the upper half contains one and not the other with room to spare.
let source = split(&ctx, [200, 200, 200], [30, 30, 30]);
let stack = brightened(MaskSource::luminance_range(0.5, 1.0, 0.15));
let pixels = render(&ctx, &source, &stack, SIZE);
let (bright, dark) = halves(&pixels);
assert!(
bright > 230,
"the bright half is inside the band and should have been lifted, got {bright}"
);
assert!(
dark < 45,
"the dark half is outside the band and must be untouched, got {dark}"
);
}
/// TRACES: FR-DEV-10
/// The property the develop view and the export path share. The mask array is
/// rasterised at a proxy size in both, but nothing in this pass may *depend*
/// on that size — a range is a function of the photograph's values, and those
/// do not change when somebody asks for a bigger picture.
#[test]
fn a_tone_band_is_the_same_mask_at_any_raster_size() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
let source = split(&ctx, [200, 200, 200], [30, 30, 30]);
let stack = brightened(MaskSource::luminance_range(0.5, 1.0, 0.15));
// A quarter of the source and twice it: one mask texel averaging sixteen
// source texels, and one source texel spread over four mask texels.
let small = halves(&render(&ctx, &source, &stack, SIZE / 4));
let large = halves(&render(&ctx, &source, &stack, SIZE * 2));
// A tolerance rather than equality: the two rasters land their edges on
// different grids, and the assertion is that the *selection* is the same,
// not that two resamplings of it are bit-identical.
assert!(
small.0.abs_diff(large.0) <= 4 && small.1.abs_diff(large.1) <= 4,
"the same band selected differently at two raster sizes: {small:?} vs {large:?}"
);
}
/// TRACES: FR-DEV-10
#[test]
fn a_colour_band_follows_hue_rather_than_brightness() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
// Short of saturation on purpose. A fully clipped channel carries no
// colour at all, and the pass pulls such a pixel back to neutral before
// the band ever sees it — which is correct, and would make this test about
// that instead.
let source = split(&ctx, [200, 40, 40], [40, 40, 200]);
// A narrow arc at red, above the chroma floor that keeps the greys out.
let stack = brightened(MaskSource::colour_range(0.0, 0.05, 0.15, 1.0, 0.05));
let pixels = render(&ctx, &source, &stack, SIZE);
let (red, blue) = halves(&pixels);
assert!(
red > 230,
"the red half is inside the arc and should have been lifted, got {red}"
);
assert!(
blue < 60,
"the blue half is a third of the circle away and must be untouched, got {blue}"
);
}
/// TRACES: FR-DEV-10
/// The greys a hue arc would otherwise sweep up.
///
/// A nearly-neutral pixel still has a hue — three almost-equal channels round
/// to one — so an arc without a chroma floor selects a haze of noise across
/// every desaturated part of the picture. It is the failure that looks like the
/// mask working badly rather than like a control that is missing.
#[test]
fn a_colour_band_ignores_the_greys() {
let Some(ctx) = ctx() else {
eprintln!("no adapter; skipping");
return;
};
// Both halves neutral, at the two brightnesses the tone test uses, so the
// only reason either could be selected is the hue arc reaching them.
let source = split(&ctx, [200, 200, 200], [30, 30, 30]);
let stack = brightened(MaskSource::colour_range(0.0, 0.5, 0.15, 1.0, 0.05));
let pixels = render(&ctx, &source, &stack, SIZE);
let (light, dark) = halves(&pixels);
assert!(
light < 215 && dark < 45,
"a grey frame was selected by a colour mask: {light}, {dark}"
);
}
+1 -1
View File
@@ -1,4 +1,4 @@
//! TRACES: FR-DSP-5
//! TRACES: FR-DSP-5 | R5
//! Zooming to 1:1 samples the source, pixel for pixel.
//!
//! FR-DSP-5: *"Fit, 1:1, and arbitrary zoom levels. At 1:1 and above, the
+44
View File
@@ -0,0 +1,44 @@
[package]
name = "dr-inference-engine"
version.workspace = true
edition.workspace = true
rust-version.workspace = true
license.workspace = true
# The one crate that names a runtime, a provider, a vendor library or a
# device (docs/inference.md §8). `dr-face` and `dr-segment` ask it for a
# session by role and never see which of these answered.
[dependencies]
thiserror.workspace = true
log.workspace = true
serde.workspace = true
serde_json.workspace = true
# `ort` is the API; what supplies it is decided once per process (§3):
# `libonnxruntime` found on disk, or `tract`. Both are behind
# `alternative-backend`, so nothing here links C on any target.
ort = { workspace = true }
ort-tract = { workspace = true, optional = true }
# dlopen, and the C types of the table it fetches. Both pure Rust;
# `libloading` is already in the tree through wgpu.
libloading = { version = "0.8", optional = true }
ort-sys = { version = "2.0.0-rc.13", default-features = false, features = ["disable-linking"], optional = true }
# The NVIDIA rungs exist on the desktop only. These features add `ort`'s
# option builders and nothing else — no linking under `alternative-backend` —
# but an Android binary has no business carrying even the option names, and
# the packaging must never be tempted to (§2, §3.1).
[target.'cfg(not(target_os = "android"))'.dependencies]
ort = { workspace = true, features = ["cuda", "tensorrt"] }
[target.'cfg(target_os = "android")'.dependencies]
ort = { workspace = true, features = ["qnn"] }
[features]
# The floor: `tract` supplies the API table when no runtime file is found, or
# always, in a build without `native`. Tests want this and nothing else.
default = ["tract"]
tract = ["dep:ort-tract"]
# Look for `libonnxruntime` on disk and hand its table to `ort`.
native = ["dep:libloading", "dep:ort-sys"]
+169
View File
@@ -0,0 +1,169 @@
//! The API table `ort` runs on, chosen once (docs/inference.md §3).
//!
//! `ort` with `alternative-backend` links no runtime and asks, on first use,
//! for an `OrtApi` — a struct of function pointers. Two things can fill it:
//! a `libonnxruntime` this module `dlopen`s, or `ort-tract`. The Rust build
//! is identical either way; the difference is whether a file was found.
use std::path::PathBuf;
use std::sync::OnceLock;
/// What supplied the table.
#[derive(Clone, Debug, PartialEq, Eq)]
pub enum Runtime {
/// Pure Rust, one core, every operator these graphs use. The floor.
Tract,
/// The C++ ONNX Runtime, loaded from `path`.
OnnxRuntime { path: PathBuf, version: String },
}
impl Runtime {
pub fn label(&self) -> String {
match self {
Runtime::Tract => "tract".into(),
Runtime::OnnxRuntime { version, .. } => format!("ONNX Runtime {version}"),
}
}
pub fn is_native(&self) -> bool {
matches!(self, Runtime::OnnxRuntime { .. })
}
}
static RUNTIME: OnceLock<Runtime> = OnceLock::new();
/// The runtime in use; tract until something installs another.
pub fn runtime() -> Runtime {
RUNTIME.get().cloned().unwrap_or(Runtime::Tract)
}
/// Install a table if none is installed yet — tract, since no directories
/// were named. What a test or an example gets, unless `DARKROOM_ORT_DIR`
/// names a runtime: the same variable the desktop honours, so an example
/// can be pointed at the runtime the app uses without learning `init`.
pub fn ensure_installed() {
if RUNTIME.get().is_none() {
let dirs: Vec<PathBuf> = std::env::var_os("DARKROOM_ORT_DIR")
.map(PathBuf::from)
.into_iter()
.collect();
install(&dirs);
}
}
/// Look for `libonnxruntime` in `dirs`, in order, and hand `ort` the first
/// table that loads; otherwise tract. Once per process.
pub fn install(dirs: &[PathBuf]) -> Runtime {
RUNTIME
.get_or_init(|| {
#[cfg(feature = "native")]
for dir in dirs {
match load_native(dir) {
Ok(rt) => return rt,
Err(e) => log::info!("inference: no runtime in {}: {e}", dir.display()),
}
}
#[cfg(not(feature = "native"))]
let _ = dirs;
install_tract()
})
.clone()
}
#[cfg(feature = "tract")]
fn install_tract() -> Runtime {
let _ = ort::set_api(ort_tract::api());
Runtime::Tract
}
#[cfg(not(feature = "tract"))]
fn install_tract() -> Runtime {
// A build with neither tract nor a runtime file has nothing to run
// models on; every `open` will report the un-set API rather than panic
// somewhere deeper.
log::error!("inference: no ONNX Runtime found and tract is not compiled in");
Runtime::Tract
}
#[cfg(feature = "native")]
fn load_native(dir: &std::path::Path) -> Result<Runtime, String> {
let name = if cfg!(target_os = "windows") {
"onnxruntime.dll"
} else if cfg!(any(target_os = "macos", target_os = "ios")) {
"libonnxruntime.dylib"
} else {
"libonnxruntime.so"
};
// An empty dir means the bare name: the system loader's search, which on
// Android includes the APK's own native libraries.
let path = if dir.as_os_str().is_empty() {
PathBuf::from(name)
} else {
find_library(dir, name).ok_or("not present")?
};
// SAFETY: the library's initialisers are ONNX Runtime's own; the symbol
// is the documented entry point with the documented signature; the table
// is copied out and the library handle is leaked, so every pointer in
// the copy stays valid for the life of the process.
unsafe {
let lib = libloading::Library::new(&path).map_err(|e| e.to_string())?;
let get_base: libloading::Symbol<
unsafe extern "system" fn() -> *const ort_sys::OrtApiBase,
> = lib.get(b"OrtGetApiBase\0").map_err(|e| e.to_string())?;
let base = get_base();
if base.is_null() {
return Err("OrtGetApiBase returned null".into());
}
let version = std::ffi::CStr::from_ptr(((*base).GetVersionString)())
.to_string_lossy()
.into_owned();
let api = ((*base).GetApi)(ort_sys::ORT_API_VERSION);
if api.is_null() {
return Err(format!(
"ONNX Runtime {version} is older than API version {}",
ort_sys::ORT_API_VERSION
));
}
if !ort::set_api((*api).clone()) {
return Err("an API table was already installed".into());
}
std::mem::forget(lib);
// Qualcomm's DSP loader finds the Hexagon skel through this variable,
// and only through it; the runtime's own directory is where the APK
// put it. Harmless anywhere else.
#[cfg(target_os = "android")]
if !dir.as_os_str().is_empty() {
std::env::set_var("ADSP_LIBRARY_PATH", dir);
}
log::info!("inference: ONNX Runtime {version} from {}", path.display());
Ok(Runtime::OnnxRuntime { path, version })
}
}
/// `libonnxruntime.so` in `dir`, or a versioned spelling of it —
/// `libonnxruntime.so.1.30.0` is what the Python wheel ships, and a package
/// that installs only the versioned file is not wrong.
#[cfg(feature = "native")]
fn find_library(dir: &std::path::Path, name: &str) -> Option<PathBuf> {
let exact = dir.join(name);
if exact.is_file() {
return Some(exact);
}
let prefix = format!("{name}.");
let mut versioned: Vec<PathBuf> = std::fs::read_dir(dir)
.ok()?
.filter_map(|e| e.ok())
.map(|e| e.path())
.filter(|p| {
p.is_file()
&& p.file_name()
.and_then(|n| n.to_str())
.is_some_and(|n| n.starts_with(&prefix))
})
.collect();
versioned.sort();
versioned.pop()
}
+121
View File
@@ -0,0 +1,121 @@
//! Compiled engines: what a rung builds once per device, and the thread that
//! builds them before anyone asks (docs/inference.md §5, §6).
//!
//! TensorRT keeps its own engine cache keyed by graph hash; QNN writes a
//! context model. Both are opaque to this crate, which tracks only *that* a
//! model compiled — by the hash of its bytes — so [`crate::open`] can tell a
//! request whether to expect the rung or its fallback.
use std::path::PathBuf;
use crate::{state, Config, Form, Rung};
enum Source {
File(PathBuf),
Bytes(&'static [u8]),
}
/// 64-bit FNV-1a. A cache key, not a checksum: two model files that collide
/// here would have to also be the same size and the same role, and the cost
/// of that is a rebuilt engine.
pub fn hash(bytes: &[u8]) -> u64 {
let mut h = 0xcbf2_9ce4_8422_2325u64;
for &b in bytes {
h ^= b as u64;
h = h.wrapping_mul(0x0000_0100_0000_01b3);
}
h
}
/// The cache entry for `bytes` compiled on `rung`.
pub fn key(rung: Rung, bytes: &[u8]) -> String {
key_of(rung, hash(bytes))
}
/// The same, from a hash already taken.
pub fn key_of(rung: Rung, hash: u64) -> String {
format!("{}:{:016x}", rung.label(), hash)
}
/// Where QNN's compiled context for `bytes` lives.
pub fn context_path(cfg: &Config, bytes: &[u8]) -> PathBuf {
cfg.cache_dir
.join("qnn")
.join(format!("{:016x}_ctx.onnx", hash(bytes)))
}
/// After the probe: compile every configured model the selected rung can
/// take, smallest first, recording each as it lands.
pub fn run() {
let (rung, cfg) = {
let s = state().lock().unwrap();
(crate::current_rung(&s), s.config.clone())
};
if !rung.compiles() {
return;
}
// Smallest first, so the detector — the one that runs per image — is
// ready soonest (§6 step 3).
let mut jobs: Vec<(crate::Role, Source, u64)> = cfg
.models
.iter()
.filter(|(role, _)| rung.serves(*role))
.filter_map(|(role, path)| {
let (path, form) = crate::resolve_model(*role, path);
(form == rung.form(*role)).then(|| {
let size = std::fs::metadata(&path).map(|m| m.len()).unwrap_or(0);
(*role, Source::File(path), size)
})
})
.chain(cfg.embedded.iter().filter_map(|(role, bytes)| {
// An embedded model has no int8 sibling to offer a rung that
// wants one; it runs on that rung's fallback.
(rung.serves(*role) && rung.form(*role) == Form::F32).then_some((
*role,
Source::Bytes(bytes),
bytes.len() as u64,
))
}))
.collect();
jobs.sort_by_key(|j| j.2);
state().lock().unwrap().wanted = jobs.len();
for (role, source, _) in jobs {
let (bytes, name) = match &source {
Source::File(path) => match std::fs::read(path) {
Ok(b) => (b, path.display().to_string()),
Err(_) => continue,
},
Source::Bytes(b) => (b.to_vec(), format!("embedded {role:?}")),
};
let key = key(rung, &bytes);
if state().lock().unwrap().cache.compiled.contains(&key) {
continue;
}
log::info!("inference: compiling {name} for {}", rung.label());
let started = std::time::Instant::now();
match crate::session::build(rung, role, &bytes, &cfg) {
Ok(session) => {
drop(session);
let mut s = state().lock().unwrap();
s.cache.compiled.insert(key);
crate::probe::write_cache(&s.config, &s.cache);
log::info!(
"inference: {name} ready on {} in {:.1} s",
rung.label(),
started.elapsed().as_secs_f64()
);
}
Err(e) => {
// This model stays on the fallback; the others still get
// their engine. A corrected model file changes the hash and
// is retried.
log::warn!(
"inference: {name} will not compile for {}: {e}",
rung.label()
);
}
}
}
}
+623
View File
@@ -0,0 +1,623 @@
//! Which runtime, which provider and which model form — decided once per
//! device, and the only crate that knows the answer (docs/inference.md).
//!
//! Consumers ask for a session by [`Role`] and get `ort`'s `Session` back;
//! what built it — tract on one core, ONNX Runtime's CPU pool, a TensorRT
//! engine, the Hexagon — is this crate's business and shows up in
//! [`status`] for the settings row and nowhere else.
//!
//! The shape follows §3 of the spec: `ort` links nothing (`alternative-backend`),
//! and the first call hands it an API table from either a `libonnxruntime`
//! found on disk or from `tract`. That choice is once per process, because
//! `ort::set_api` is; everything after it — which provider, whether an engine
//! has been compiled yet — is per session and may change between two calls.
use std::collections::{BTreeSet, HashMap};
use std::path::{Path, PathBuf};
use std::sync::{Arc, Mutex, MutexGuard, OnceLock};
use std::time::{Duration, Instant};
use serde::{Deserialize, Serialize};
mod api;
mod engines;
mod probe;
mod session;
pub use api::Runtime;
pub use ort::session::Session;
/// What a model is for. The role fixes the precision rule (§7): an embedder
/// runs in f32 on every rung, a detector may run in fp16 or int8.
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, Serialize, Deserialize)]
pub enum Role {
Detector,
Embedder,
Segmenter,
Scene,
/// The dense landmark model behind the eye reading (docs/faces.md §7c).
Landmarks,
/// The eye-state and sunglasses classifiers, a few hundred kilobytes.
EyeClassifier,
/// XFeat, the panorama keypoint detector (docs/panorama.md).
Keypoints,
/// MI-GAN, the panorama border filler (docs/panorama.md §12). Plain
/// convolutions, so any rung serves it; fp16 on TensorRT and int8 on
/// the Hexagon are the point of it.
Inpainter,
}
/// Which numeric form of a model a session was built from.
///
/// `Int8` is a different network from `F32` for a detector — it finds a
/// different set of faces — which is why [`form_suffix`] exists and why a
/// caller appends it to `model_id`.
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, Serialize, Deserialize)]
pub enum Form {
F32,
Int8,
}
/// A rung of the ladder (§2). Ordered: a user override names the highest rung
/// the probe may take, and a compiling rung falls back to the one below it
/// until its engine exists.
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, PartialOrd, Ord, Serialize, Deserialize)]
pub enum Rung {
/// ONNX Runtime's CPU provider, or tract when no runtime file was found.
Cpu,
/// NVIDIA, through the CUDA provider. Desktop only.
Cuda,
/// NVIDIA, through a TensorRT engine compiled on this device. Desktop only.
TensorRt,
/// Qualcomm's Hexagon NPU through QNN, int8 models only. Android only.
Hexagon,
}
impl Rung {
pub fn label(self) -> &'static str {
match self {
Rung::Cpu => "CPU",
Rung::Cuda => "CUDA",
Rung::TensorRt => "TensorRT",
Rung::Hexagon => "Hexagon NPU",
}
}
/// The rung a request lands on while this one's engine is still being
/// compiled (§6 step 2).
fn fallback(self) -> Rung {
match self {
Rung::TensorRt => Rung::Cuda,
Rung::Hexagon | Rung::Cuda | Rung::Cpu => Rung::Cpu,
}
}
/// Whether a session on this rung needs an engine built first.
fn compiles(self) -> bool {
matches!(self, Rung::TensorRt | Rung::Hexagon)
}
/// The model form this rung wants for a role.
fn form(self, _role: Role) -> Form {
match self {
Rung::Hexagon => Form::Int8,
_ => Form::F32,
}
}
/// Whether this rung runs `role` at all. The Hexagon takes int8 graphs
/// only, and the embedder is never int8 (§7) — it runs on the CPU
/// beside a detector on the NPU, so its vectors compare across devices.
fn serves(self, role: Role) -> bool {
match self {
Rung::Hexagon => role != Role::Embedder,
_ => true,
}
}
}
/// How long a session outlives its last use unless [`Config::decay`] says
/// otherwise: long enough for the next click, short enough that a session's
/// GPU or NPU memory does not sit under the develop view for long.
pub const DEFAULT_DECAY: Duration = Duration::from_secs(30);
/// What [`init`] is told once, at launch.
#[derive(Clone, Debug, Default)]
pub struct Config {
/// Where to look for `libonnxruntime`, in order. An empty path means "the
/// bare library name through the system loader", which is how the APK's
/// own copy is found on Android.
pub runtime_dirs: Vec<PathBuf>,
/// Probe cache and compiled engines (§4, §5). Disposable.
pub cache_dir: PathBuf,
/// The canonical model files on this device, so engines can be compiled
/// ahead of the first request for them.
pub models: Vec<(Role, PathBuf)>,
/// Models compiled into the binary, for the same reason.
pub embedded: Vec<(Role, &'static [u8])>,
/// The highest rung the user allows; `None` is "the best that works".
pub ceiling: Option<Rung>,
/// ONNX Runtime's intra-op pool; 0 picks from the core count.
pub threads: usize,
/// How long an unused session stays loaded. Zero means the default.
pub decay: Duration,
}
/// One line for the settings row, and the numbers behind the progress row.
#[derive(Clone, Debug)]
pub struct Status {
pub runtime: Runtime,
/// The rung selected, or the floor while the probe is still running.
pub rung: Rung,
/// Why — "probe passed", or the failure that demoted the rung above.
pub reason: String,
pub probing: bool,
/// Engines compiled and engines wanted, for a compiling rung; `(0, 0)`
/// otherwise.
pub engines: (usize, usize),
/// Every rung above the selected one that was tried, and why it lost.
pub failed: Vec<(Rung, String)>,
}
impl Status {
/// "Hexagon NPU · int8 · ONNX Runtime 1.29" — the settings row's text.
pub fn line(&self) -> String {
let form = match self.rung {
Rung::Hexagon => " · int8",
Rung::TensorRt => " · fp16",
_ => "",
};
format!("{}{} · {}", self.rung.label(), form, self.runtime.label())
}
}
/// A model the caller can run, whatever is or is not loaded right now.
///
/// Holds the bytes, not a session. [`Model::acquire`] finds the loaded copy
/// in the registry — shared with every other holder of the same model —
/// or loads one, and every acquire refreshes the copy's last-used time.
/// The reaper unloads anything idle for [`Config::decay`]; a scan that runs
/// the detector on every image never lets it go idle, a click in the
/// develop view lets the segmenter go after a quiet spell, and a handle
/// used again after that simply loads again. Nobody states a policy.
///
/// The registry key includes the rung, so a reload after a compiled engine
/// has landed moves up to it by itself (§6 step 4).
pub struct Model {
role: Role,
form: Form,
bytes: Arc<[u8]>,
/// `engines::hash` of the bytes, taken once: an acquire per tile of a
/// border fill must not hash 28 MB each time.
hash: u64,
}
/// A loaded session, held for one `run` and its output decoding.
pub struct Acquired {
entry: Arc<Loaded>,
}
struct Loaded {
rung: Rung,
session: Mutex<Session>,
last_used: Mutex<Instant>,
}
impl Model {
/// The loaded session, loading it if the reaper took it. Lock it for
/// one run; a scan and a develop click can want the same detector at
/// once, and the second waits on the first.
pub fn acquire(&self) -> Result<Acquired, Error> {
acquire(self.role, self.form, &self.bytes, self.hash)
}
pub fn form(&self) -> Form {
self.form
}
}
impl Acquired {
pub fn lock(&self) -> MutexGuard<'_, Session> {
self.entry.session.lock().unwrap_or_else(|e| e.into_inner())
}
/// Where this session runs.
pub fn rung(&self) -> Rung {
self.entry.rung
}
}
impl Drop for Acquired {
fn drop(&mut self) {
// The clock starts when the use ends, not when it began: a long run
// is not idle time.
*self.entry.last_used.lock().unwrap() = Instant::now();
}
}
type Registry = HashMap<String, Arc<Loaded>>;
static REGISTRY: OnceLock<Mutex<Registry>> = OnceLock::new();
fn registry() -> &'static Mutex<Registry> {
REGISTRY.get_or_init(|| {
std::thread::Builder::new()
.name("inference-reaper".into())
.spawn(|| loop {
std::thread::sleep(Duration::from_secs(5));
release_idle();
})
.expect("spawn inference reaper");
Mutex::new(HashMap::new())
})
}
fn acquire(role: Role, form: Form, bytes: &Arc<[u8]>, hash: u64) -> Result<Acquired, Error> {
api::ensure_installed();
let (rung, cfg) = {
let s = state().lock().unwrap();
let selected = current_rung(&s);
(
effective_rung(&s, selected, role, form, hash),
s.config.clone(),
)
};
let key = format!("{role:?}:{}", engines::key_of(rung, hash));
if let Some(entry) = registry().lock().unwrap().get(&key).cloned() {
*entry.last_used.lock().unwrap() = Instant::now();
return Ok(Acquired { entry });
}
// Built outside the registry lock: a TensorRT engine load is long enough
// that another role's acquire should not wait on it.
let session = session::build(rung, role, bytes, &cfg)?;
log::debug!("inference: {role:?} loaded on {}", rung.label());
let entry = Arc::new(Loaded {
rung,
session: Mutex::new(session),
last_used: Mutex::new(Instant::now()),
});
let mut reg = registry().lock().unwrap();
// Two acquires raced; keep the first, drop this one.
let entry = reg.entry(key).or_insert_with(|| entry.clone()).clone();
Ok(Acquired { entry })
}
/// Unload every session idle for longer than the decay. The reaper does
/// this every five seconds. A session in use survives until its run ends:
/// the `Acquired` holds it, the registry merely forgets it.
pub fn release_idle() {
let decay = match state().lock().unwrap().config.decay {
Duration::ZERO => DEFAULT_DECAY,
d => d,
};
let now = Instant::now();
registry()
.lock()
.unwrap()
.retain(|_, e| now.duration_since(*e.last_used.lock().unwrap()) < decay);
}
/// Unload every session now, decay or not — what a low-memory signal
/// asks for. Sessions mid-run finish first.
pub fn release_all() {
registry().lock().unwrap().clear();
}
/// Unload every session of `role` now — "I am done segmenting".
pub fn unload(role: Role) {
let prefix = format!("{role:?}:");
registry()
.lock()
.unwrap()
.retain(|k, _| !k.starts_with(&prefix));
}
/// How many sessions are loaded, for the settings row and the tests.
pub fn loaded() -> usize {
registry().lock().unwrap().len()
}
#[derive(Debug, thiserror::Error)]
pub enum Error {
#[error(transparent)]
Inference(#[from] ort::Error),
#[error("reading model: {0}")]
Io(#[from] std::io::Error),
}
/// What the probe writes and the next launch reads (§4 step 3).
#[derive(Clone, Debug, Default, Serialize, Deserialize)]
struct Cache {
/// Runtime, driver, hardware and model identity; any change re-probes.
fingerprint: String,
rung: Option<Rung>,
reason: String,
/// Model hashes whose engine exists on disk, per compiling rung.
compiled: BTreeSet<String>,
/// Rungs that failed under this fingerprint, and why. Not retried until
/// the fingerprint changes: a wedged driver must not cost every launch
/// thirty seconds.
failed: Vec<(Rung, String)>,
}
struct State {
config: Config,
cache: Cache,
probing: bool,
wanted: usize,
}
static STATE: OnceLock<Mutex<State>> = OnceLock::new();
fn state() -> &'static Mutex<State> {
STATE.get_or_init(|| {
Mutex::new(State {
config: Config::default(),
cache: Cache::default(),
probing: false,
wanted: 0,
})
})
}
/// Choose the runtime and start the probe. Idempotent; the first call wins.
///
/// Returns at once: the probe and any engine compilation run on their own
/// low-priority thread, and every request meanwhile is served by the floor
/// (§4). Never blocks the first frame.
pub fn init(config: Config) {
let runtime = api::install(&config.runtime_dirs);
{
let mut s = state().lock().unwrap();
if s.probing || s.cache.rung.is_some() {
return;
}
s.config = config;
s.probing = true;
}
log::info!("inference: runtime {}", runtime.label());
std::thread::Builder::new()
.name("inference-probe".into())
.spawn(move || {
probe::run(runtime);
engines::run();
})
.expect("spawn inference probe");
}
/// Make sure `ort` has an API table, for code that drives `ort` directly.
/// [`open`] does this itself; only the M1 probe example needs it by name.
pub fn ensure_runtime() {
api::ensure_installed();
}
/// The line for the settings row.
pub fn status() -> Status {
let s = state().lock().unwrap();
let rung = current_rung(&s);
Status {
runtime: api::runtime(),
rung,
reason: s.cache.reason.clone(),
failed: s.cache.failed.clone(),
probing: s.probing,
engines: if rung.compiles() {
(s.cache.compiled.len(), s.wanted)
} else {
(0, 0)
},
}
}
fn current_rung(s: &State) -> Rung {
if s.probing {
Rung::Cpu
} else {
s.cache.rung.unwrap_or(Rung::Cpu)
}
}
/// The file to load for `role` under the current selection, and its form.
///
/// A rung that wants int8 gets the `.int8.onnx` sibling of the canonical file
/// if it exists; otherwise the canonical file, on the rung's fallback. A
/// caller adds [`form_suffix`] to the `model_id` it records.
pub fn resolve_model(role: Role, canonical: &Path) -> (PathBuf, Form) {
let rung = current_rung(&state().lock().unwrap());
if rung.serves(role) && rung.form(role) == Form::Int8 {
let sibling = int8_sibling(canonical);
if sibling.is_file() {
return (sibling, Form::Int8);
}
}
(canonical.to_path_buf(), Form::F32)
}
fn int8_sibling(canonical: &Path) -> PathBuf {
let stem = canonical
.file_stem()
.map(|s| s.to_string_lossy().into_owned())
.unwrap_or_default();
canonical.with_file_name(format!("{stem}.int8.onnx"))
}
/// What a form appends to a detector's `model_id` (§7).
pub fn form_suffix(form: Form) -> &'static str {
match form {
Form::F32 => "",
Form::Int8 => "_i8",
}
}
/// A handle on the model `bytes` in `role`.
///
/// Loads it once here, so a graph the runtime rejects fails at
/// construction and not on the first image; what happens to that session
/// afterwards is the registry's business (see [`Model`]).
///
/// Works without [`init`] — a test, or the examples — by installing tract
/// and using the CPU rung, which is exactly what every consumer did before
/// this crate existed.
pub fn open(role: Role, form: Form, bytes: &[u8]) -> Result<Model, Error> {
let bytes: Arc<[u8]> = Arc::from(bytes);
let hash = engines::hash(&bytes);
acquire(role, form, &bytes, hash)?;
Ok(Model {
role,
form,
bytes,
hash,
})
}
/// Where a request lands: the selected rung unless the role's precision rule,
/// the form on offer, or a missing engine says one lower (§6 step 4).
fn effective_rung(s: &State, selected: Rung, role: Role, form: Form, hash: u64) -> Rung {
let mut rung = selected;
if !rung.serves(role) || rung.form(role) != form {
// The embedder on a Hexagon device, or an f32 detector where the int8
// sibling was missing: neither can go to the NPU.
rung = rung.fallback();
}
if rung.compiles() && !s.cache.compiled.contains(&engines::key_of(rung, hash)) {
rung = rung.fallback();
}
rung
}
#[cfg(test)]
mod tests {
use super::*;
/// The registry is one per process, so these run one at a time.
static SERIAL: Mutex<()> = Mutex::new(());
fn serial() -> MutexGuard<'static, ()> {
SERIAL.lock().unwrap_or_else(|e| e.into_inner())
}
/// The smallest shipped graph, if this checkout has the weights; a test
/// suite that needs a research-licensed download is one that does not
/// run in CI (docs/faces.md §3), so absence is a skip.
fn probe_bytes() -> Option<Vec<u8>> {
let path = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../../models/face/scrfd_500m_640.onnx"
);
let bytes = std::fs::read(path).ok()?;
(bytes.len() > 100_000).then_some(bytes)
}
#[test]
fn two_handles_on_one_model_share_one_session() {
let _serial = serial();
let Some(bytes) = probe_bytes() else { return };
release_all();
let a = open(Role::Detector, Form::F32, &bytes).unwrap();
let b = open(Role::Detector, Form::F32, &bytes).unwrap();
assert_eq!(loaded(), 1);
let (x, y) = (a.acquire().unwrap(), b.acquire().unwrap());
assert!(Arc::ptr_eq(&x.entry, &y.entry));
}
#[test]
fn a_released_model_reloads_on_its_next_use() {
let _serial = serial();
let Some(bytes) = probe_bytes() else { return };
release_all();
let model = open(Role::Detector, Form::F32, &bytes).unwrap();
assert_eq!(loaded(), 1);
release_all();
assert_eq!(loaded(), 0);
let acquired = model.acquire().unwrap();
assert_eq!(loaded(), 1);
assert_eq!(acquired.lock().inputs().len(), 1);
}
#[test]
fn an_idle_session_decays_and_a_used_one_does_not() {
let _serial = serial();
let Some(bytes) = probe_bytes() else { return };
release_all();
state().lock().unwrap().config.decay = Duration::from_millis(50);
let model = open(Role::Detector, Form::F32, &bytes).unwrap();
// Used within the decay: stays.
std::thread::sleep(Duration::from_millis(30));
drop(model.acquire().unwrap());
release_idle();
assert_eq!(loaded(), 1);
// Idle past it: goes.
std::thread::sleep(Duration::from_millis(80));
release_idle();
assert_eq!(loaded(), 0);
state().lock().unwrap().config.decay = Duration::ZERO;
}
#[test]
fn unload_by_role_leaves_the_other_roles() {
let _serial = serial();
let Some(bytes) = probe_bytes() else { return };
release_all();
let _d = open(Role::Detector, Form::F32, &bytes).unwrap();
let _s = open(Role::Segmenter, Form::F32, &bytes).unwrap();
assert_eq!(loaded(), 2);
unload(Role::Segmenter);
assert_eq!(loaded(), 1);
}
#[test]
fn the_hexagon_never_takes_the_embedder() {
assert!(!Rung::Hexagon.serves(Role::Embedder));
assert!(Rung::Hexagon.serves(Role::Detector));
assert_eq!(Rung::Hexagon.form(Role::Detector), Form::Int8);
// A detector offered in f32 on a Hexagon device lands on the CPU.
let s = State {
config: Config::default(),
cache: Cache {
rung: Some(Rung::Hexagon),
..Cache::default()
},
probing: false,
wanted: 0,
};
assert_eq!(
effective_rung(
&s,
Rung::Hexagon,
Role::Embedder,
Form::F32,
engines::hash(b"")
),
Rung::Cpu
);
assert_eq!(
effective_rung(
&s,
Rung::Hexagon,
Role::Detector,
Form::F32,
engines::hash(b"")
),
Rung::Cpu
);
// An int8 detector whose context is not compiled yet: also the CPU.
assert_eq!(
effective_rung(
&s,
Rung::Hexagon,
Role::Detector,
Form::Int8,
engines::hash(b"")
),
Rung::Cpu
);
}
#[test]
fn the_status_line_reads_as_the_floor_before_init() {
let s = status();
assert_eq!(s.rung, Rung::Cpu);
assert!(s.line().starts_with("CPU"), "{}", s.line());
}
}
+302
View File
@@ -0,0 +1,302 @@
//! Walk the ladder, once, by building real sessions (docs/inference.md §4).
//!
//! A rung is taken when a session builds on it, runs, and is faster than
//! the floor. Both halves matter: a provider can register and then fail at
//! partition time, and a provider can take a graph — or quietly hand most
//! of it back to the CPU — and run it slower than the CPU would have. The outcome is cached against a fingerprint of the
//! runtime, the driver, the hardware and the models, and trusted until any
//! of those changes.
use std::path::{Path, PathBuf};
use std::time::Instant;
use crate::{api::Runtime, state, Cache, Config, Form, Role, Rung};
/// The rungs to try on this platform, best first, under the user's ceiling.
fn ladder(ceiling: Option<Rung>) -> Vec<Rung> {
#[cfg(target_os = "android")]
let all = [Rung::Hexagon];
#[cfg(not(target_os = "android"))]
let all = [Rung::TensorRt, Rung::Cuda];
all.into_iter()
.filter(|r| ceiling.is_none_or(|c| *r <= c))
.collect()
}
/// The probe body. Sets the cache and clears `probing` when done; never
/// panics out, because a failed probe is a result (the floor) and not an
/// error.
pub fn run(runtime: Runtime) {
let cfg = state().lock().unwrap().config.clone();
let fingerprint = fingerprint(&runtime, &cfg);
if let Some(cached) = read_cache(&cfg) {
if cached.fingerprint == fingerprint && cached.rung.is_some() {
log::info!(
"inference: cached selection {} ({})",
cached.rung.unwrap().label(),
cached.reason
);
finish(cached);
return;
}
}
let mut cache = Cache {
fingerprint,
..Cache::default()
};
if !runtime.is_native() {
cache.rung = Some(Rung::Cpu);
cache.reason = "no ONNX Runtime found; tract on one core".into();
write_cache(&cfg, &cache);
finish(cache);
return;
}
let Some((role, canonical)) = probe_model(&cfg) else {
cache.rung = Some(Rung::Cpu);
cache.reason = "no model to probe with".into();
write_cache(&cfg, &cache);
finish(cache);
return;
};
let floor = match time_rung(Rung::Cpu, role, &canonical, &cfg) {
Ok((ms, _)) => ms,
Err(e) => {
// The CPU provider failing is the runtime failing; there is
// nothing below it to try, and the reason is worth reading.
cache.rung = Some(Rung::Cpu);
cache.reason = format!("CPU provider failed: {e}");
write_cache(&cfg, &cache);
finish(cache);
return;
}
};
log::info!("inference: floor {floor:.1} ms on the CPU provider");
for rung in ladder(cfg.ceiling) {
match time_rung(rung, role, &canonical, &cfg) {
Ok((ms, key)) if ms < floor => {
cache.rung = Some(rung);
cache.reason = format!("{ms:.1} ms against {floor:.1} ms on the CPU");
if let Some(key) = key {
cache.compiled.insert(key);
}
break;
}
Ok((ms, _)) => {
let why = format!("{ms:.1} ms, slower than the CPU's {floor:.1} ms");
log::info!("inference: {} rejected: {why}", rung.label());
cache.failed.push((rung, why));
}
Err(e) => {
log::info!("inference: {} failed: {e}", rung.label());
cache.failed.push((rung, e));
}
}
}
if cache.rung.is_none() {
cache.rung = Some(Rung::Cpu);
cache.reason = match cache.failed.first() {
Some((r, why)) => format!("{} {}", r.label(), first_line(why)),
None => "the only rung on this platform".into(),
};
}
write_cache(&cfg, &cache);
finish(cache);
}
fn finish(cache: Cache) {
let mut s = state().lock().unwrap();
s.cache = cache;
s.probing = false;
}
/// The smallest detector, or the smallest model of any role if there is
/// none. A ~2 MB detector is the cheapest real test of a provider, and the
/// detector is the role the int8 forms exist for — the eye classifiers are
/// smaller still, and a Hexagon probed with one would fail for want of a
/// form nobody ships.
fn probe_model(cfg: &Config) -> Option<(Role, PathBuf)> {
let smallest = |want: Option<Role>| {
cfg.models
.iter()
.filter(|(role, _)| want.is_none_or(|w| *role == w))
.filter_map(|(role, path)| {
let size = std::fs::metadata(path).ok()?.len();
Some((size, *role, path.clone()))
})
.min_by_key(|(size, _, _)| *size)
.map(|(_, role, path)| (role, path))
};
smallest(Some(Role::Detector)).or_else(|| smallest(None))
}
/// Build, run once for the engine, then time three runs; the median in
/// milliseconds and, for a compiling rung, the cache key of the engine this
/// just built.
fn time_rung(
rung: Rung,
role: Role,
canonical: &Path,
cfg: &Config,
) -> Result<(f64, Option<String>), String> {
let want = rung.form(role);
let path = match want {
Form::Int8 => {
let p = crate::int8_sibling(canonical);
if !p.is_file() {
return Err(format!("no int8 form of {}", canonical.display()));
}
p
}
Form::F32 => canonical.to_path_buf(),
};
let bytes = std::fs::read(&path).map_err(|e| e.to_string())?;
let started = Instant::now();
let mut session =
crate::session::build(rung, role, &bytes, cfg).map_err(|e| first_line(&e.to_string()))?;
log::info!(
"inference: {} session built in {:.1} s",
rung.label(),
started.elapsed().as_secs_f64()
);
let shape: Vec<usize> = session.inputs()[0]
.dtype()
.tensor_shape()
.ok_or("model input is not a tensor")?
.iter()
.map(|&d| if d > 0 { d as usize } else { 1 })
.collect();
let zeros = vec![0f32; shape.iter().product()];
let run = |session: &mut ort::session::Session| -> Result<f64, String> {
let input = ort::value::Tensor::from_array((shape.clone(), zeros.clone()))
.map_err(|e| e.to_string())?;
let t = Instant::now();
let out = session
.run(ort::inputs![input])
.map_err(|e| e.to_string())?;
let _ = out[0]
.try_extract_tensor::<f32>()
.map_err(|e| e.to_string())?;
Ok(t.elapsed().as_secs_f64() * 1e3)
};
run(&mut session)?;
let mut times = [run(&mut session)?, run(&mut session)?, run(&mut session)?];
times.sort_by(|a, b| a.partial_cmp(b).unwrap());
let key = rung.compiles().then(|| crate::engines::key(rung, &bytes));
Ok((times[1], key))
}
/// The part of a provider's error a person can act on. ONNX Runtime's
/// begin with a source path and a C++ template signature; the words —
/// "CUDA failure 999: unknown error", "FAIL : Failed to load library" —
/// come after, and the settings row has room for one line of them.
fn first_line(s: &str) -> String {
let line = s.lines().next().unwrap_or("");
let start = ["failure", "FAIL :", "Error:", "error:"]
.iter()
.filter_map(|m| line.find(m))
.min()
.unwrap_or(0);
line[start..].chars().take(200).collect()
}
/// Everything a change of which should re-probe: the runtime and where it
/// came from, this crate, the platform, the driver or SoC, and the models.
fn fingerprint(runtime: &Runtime, cfg: &Config) -> String {
let mut parts = vec![
format!("engine {}", env!("CARGO_PKG_VERSION")),
format!("{} {}", std::env::consts::OS, std::env::consts::ARCH),
match runtime {
Runtime::Tract => "tract".to_string(),
Runtime::OnnxRuntime { path, version } => format!("ort {version} {}", path.display()),
},
device_identity(),
];
for (role, bytes) in &cfg.embedded {
parts.push(format!(
"{role:?} embedded {:016x}",
crate::engines::hash(bytes)
));
}
for (role, path) in &cfg.models {
let hash = std::fs::read(path)
.map(|b| crate::engines::hash(&b))
.unwrap_or(0);
parts.push(format!("{role:?} {hash:016x}"));
let int8 = crate::int8_sibling(path);
if let Ok(b) = std::fs::read(&int8) {
parts.push(format!("{role:?} int8 {:016x}", crate::engines::hash(&b)));
}
}
parts.join("\n")
}
#[cfg(target_os = "linux")]
fn device_identity() -> String {
// The NVIDIA driver's version line; absent means no NVIDIA driver.
std::fs::read_to_string("/proc/driver/nvidia/version")
.ok()
.and_then(|s| s.lines().next().map(str::to_string))
.unwrap_or_else(|| "no nvidia driver".into())
}
#[cfg(target_os = "android")]
fn device_identity() -> String {
// The SoC and the vendor's build: a Hexagon appears or disappears with
// either.
format!(
"{} {}",
system_property("ro.soc.model"),
system_property("ro.build.version.incremental")
)
}
#[cfg(target_os = "android")]
fn system_property(name: &str) -> String {
extern "C" {
fn __system_property_get(
name: *const std::ffi::c_char,
value: *mut std::ffi::c_char,
) -> i32;
}
let name = std::ffi::CString::new(name).unwrap();
let mut buf = [0u8; 92]; // PROP_VALUE_MAX
// SAFETY: bionic's documented call; the buffer is PROP_VALUE_MAX bytes.
let n = unsafe { __system_property_get(name.as_ptr(), buf.as_mut_ptr().cast()) };
String::from_utf8_lossy(&buf[..n.max(0) as usize]).into_owned()
}
#[cfg(not(any(target_os = "linux", target_os = "android")))]
fn device_identity() -> String {
String::new()
}
fn cache_path(cfg: &Config) -> PathBuf {
cfg.cache_dir.join("backend.json")
}
fn read_cache(cfg: &Config) -> Option<Cache> {
let text = std::fs::read_to_string(cache_path(cfg)).ok()?;
serde_json::from_str(&text).ok()
}
/// Written whole and renamed into place, so a reader never sees half.
pub fn write_cache(cfg: &Config, cache: &Cache) {
if cfg.cache_dir.as_os_str().is_empty() {
return;
}
let path = cache_path(cfg);
let tmp = path.with_extension("json.tmp");
let _ = std::fs::create_dir_all(&cfg.cache_dir);
if let Ok(text) = serde_json::to_string_pretty(cache) {
if std::fs::write(&tmp, text).is_ok() {
let _ = std::fs::rename(&tmp, &path);
}
}
}
+125
View File
@@ -0,0 +1,125 @@
//! One session builder per rung (docs/inference.md §2, §7, §9).
use ort::session::Session;
use crate::{Config, Role, Rung};
/// Build a session for `bytes` on `rung`.
///
/// Not strict about the CPU: `session.disable_cpu_ep_fallback` was tried as
/// the probe's proof that a provider took the graph, and it refuses the
/// Hexagon over the ten quantise/dequantise nodes at the graph's edges that
/// QNN declines by policy and that cost microseconds. The probe's proof is
/// its clock instead (§4): a provider that hands real work to the CPU is
/// slower than the CPU floor and rejected by the same measurement.
pub fn build(rung: Rung, role: Role, bytes: &[u8], cfg: &Config) -> ort::Result<Session> {
// No optimisation level named. ONNX Runtime's default is already its
// fullest, and on tract any level but "disabled" means `into_optimized`,
// whose optimiser divides by zero inside yolo26n-seg (tract-data
// `stack_tensors`) — a panic across the C API, which is an abort. The
// app never asked tract for that and does not start now.
let mut b = Session::builder()?.with_intra_threads(threads(cfg))?;
// A Hexagon session loads the compiled context when there is one and
// compiles it from the model when there is not; the engine thread is
// what makes the second case rare (§6).
let context = (rung == Rung::Hexagon).then(|| crate::engines::context_path(cfg, bytes));
let ready = context.as_ref().is_some_and(|p| p.is_file());
b = providers(
b,
rung,
role,
cfg,
if ready { None } else { context.as_deref() },
)?;
match (ready, context) {
(true, Some(path)) => b.commit_from_file(path),
_ => b.commit_from_memory(bytes),
}
}
/// The intra-op pool: what the config says, else the cores less two for
/// the compositor and the decoder (§9). tract ignores it.
fn threads(cfg: &Config) -> usize {
if cfg.threads > 0 {
return cfg.threads;
}
std::thread::available_parallelism()
.map(|n| n.get().saturating_sub(2).max(1))
.unwrap_or(1)
}
#[cfg(not(target_os = "android"))]
fn providers(
b: ort::session::builder::SessionBuilder,
rung: Rung,
role: Role,
cfg: &Config,
_generate_context: Option<&std::path::Path>,
) -> ort::Result<ort::session::builder::SessionBuilder> {
use ort::ep;
match rung {
Rung::Cpu => Ok(b),
Rung::Cuda => {
Ok(b.with_execution_providers([ep::CUDA::default().build().error_on_failure()])?)
}
Rung::TensorRt => {
let cache = cfg.cache_dir.join("tensorrt");
let _ = std::fs::create_dir_all(&cache);
let cache = cache.to_string_lossy().into_owned();
// fp16 for everything but the embedder, whose comparability
// across devices is worth more than its 0.2 ms (§7). The
// workspace cap keeps the develop view's tiles on the card
// (NFR-RES-2). CUDA behind it takes any node TensorRT declines.
Ok(b.with_execution_providers([
ep::TensorRT::default()
.with_fp16(role != Role::Embedder)
.with_engine_cache(true)
.with_engine_cache_path(&cache)
.with_timing_cache(true)
.with_timing_cache_path(&cache)
.with_max_workspace_size(512 << 20)
.build()
.error_on_failure(),
ep::CUDA::default().build(),
])?)
}
Rung::Hexagon => unreachable!("the Hexagon rung is not on a desktop ladder"),
}
}
#[cfg(target_os = "android")]
fn providers(
b: ort::session::builder::SessionBuilder,
rung: Rung,
_role: Role,
_cfg: &Config,
generate_context: Option<&std::path::Path>,
) -> ort::Result<ort::session::builder::SessionBuilder> {
use ort::ep;
match rung {
Rung::Cpu => Ok(b),
Rung::Hexagon => {
// The HTP compiles the graph once per device (0.8–1.7 s here).
// With `ep.context_enable` ONNX Runtime writes the compiled
// context beside the probe cache; the next session loads that
// file as its model and skips the compile (§5).
let mut b = b;
if let Some(ctx) = generate_context {
let _ = std::fs::create_dir_all(ctx.parent().unwrap());
b = b
.with_config_entry("ep.context_enable", "1")?
.with_config_entry("ep.context_file_path", ctx.to_string_lossy())?
.with_config_entry("ep.context_embed_mode", "0")?;
}
// Quantise/dequantise at the graph's edges stay on the NPU too,
// so a strict build is a whole-graph build.
Ok(b.with_execution_providers([ep::QNN::default()
.with_backend_path("libQnnHtp.so")
.with_performance_mode(ep::qnn::PerformanceMode::Burst)
.with_offload_graph_io_quantization(false)
.build()
.error_on_failure()])?)
}
Rung::Cuda | Rung::TensorRt => unreachable!("no NVIDIA rung on Android"),
}
}
+38
View File
@@ -0,0 +1,38 @@
[package]
name = "dr-pano"
version.workspace = true
edition.workspace = true
rust-version.workspace = true
license.workspace = true
# Guards against a Git LFS pointer being embedded in place of the weights.
build = "build.rs"
[dependencies]
thiserror.workspace = true
log.workspace = true
# Inference for the learned keypoint detector, on the same footing as
# `dr-segment`: `ort` is the API, `dr-inference-engine` decides what runs
# it (docs/inference.md), and both are optional so that the geometry —
# matching, the rotation solve, the projections — is a dependency-free crate
# that tests without a model.
ort = { workspace = true, optional = true }
dr-inference-engine = { workspace = true, optional = true }
ndarray = { workspace = true, optional = true }
[dev-dependencies]
# The example aligns real frames from their embedded previews.
dr-decode.workspace = true
dr-types.workspace = true
env_logger.workspace = true
[features]
default = ["xfeat", "embedded-model"]
# The XFeat detector (FR-MRG-8) and the MI-GAN filler (FR-MRG-4). Off, the
# crate has no model and no runtime — a build that only wants the geometry.
xfeat = ["dep:ort", "dep:dr-inference-engine", "dep:ndarray"]
# Compile the weights into the binary, for the same reason `dr-segment` does:
# Android hands the app no path to read a model from (ARCH §6.9).
embedded-model = ["xfeat"]
+48
View File
@@ -0,0 +1,48 @@
//! Check the model is a model and not an LFS pointer.
//!
//! `models/keypoints/*.onnx` is stored in Git LFS (see `.gitattributes`). A
//! clone made without git-lfs, or with `GIT_LFS_SKIP_SMUDGE` set, leaves a
//! ~130-byte text pointer at that path instead of the weights, and
//! `include_bytes!` would embed it without complaint. Same guard as
//! `dr-segment`'s, for the same failure.
use std::path::Path;
const MODELS: &[&str] = &[
"../../models/keypoints/xfeat-1024.onnx",
"../../models/keypoints/xfeat-768.onnx",
];
fn main() {
for m in MODELS {
println!("cargo:rerun-if-changed={m}");
}
println!("cargo:rerun-if-changed=build.rs");
if std::env::var_os("CARGO_FEATURE_EMBEDDED_MODEL").is_none() {
return;
}
for model in MODELS.iter().copied() {
check(model);
}
}
fn check(model: &str) {
let path = Path::new(model);
let Ok(bytes) = std::fs::read(path) else {
panic!(
"\n\n{model} is missing.\n\
It ships in Git LFS. Run `git lfs install && git lfs pull`, or build \
with `--no-default-features` for a geometry-only build.\n"
);
};
if bytes.starts_with(b"version https://git-lfs.github.com/spec/") {
panic!(
"\n\n{model} is a Git LFS pointer, not the model.\n\
Run `git lfs install && git lfs pull`, or build with \
`--no-default-features` for a geometry-only build.\n"
);
}
}
+190
View File
@@ -0,0 +1,190 @@
//! Align real frames from their embedded previews and draw the result.
//!
//! ```sh
//! cargo run -p dr-pano --example align --release -- fixtures/pano/2025-08-05/*.CR2
//! cargo run -p dr-pano --example align --release -- out-prefix frame1.CR2 frame2.CR2 …
//! ```
//!
//! The point of looking rather than asserting: a rotation solve that is
//! numerically converged and geometrically wrong — a mirrored axis, a
//! transposed homography, an orientation applied the wrong way — produces
//! perfectly plausible residuals and a picture that is obviously broken.
//! This writes `<prefix>-cyl.ppm`: every frame's preview warped onto a
//! cylinder and averaged where they overlap, at a size that fits on a
//! screen. Ghosting in the overlaps is the alignment error, made visible.
//!
//! Previews, not RAW: the alignment runs on proxies in the application too
//! (FR-MRG-7), and a camera's embedded JPEG is a proxy the decoder already
//! extracts in milliseconds. What is different from the real path is only
//! that the pixels are the camera's rendering rather than ours, which the
//! geometry does not care about.
use std::path::PathBuf;
use std::time::Instant;
use dr_pano::bundle::Cameras;
use dr_pano::{align, xfeat::XFeat, AlignOptions, Gray, Projection};
fn main() {
env_logger::init();
let mut args: Vec<String> = std::env::args().skip(1).collect();
if args.is_empty() {
eprintln!("usage: align [out-prefix] <frame>...");
std::process::exit(2);
}
let prefix =
if args[0].ends_with(".CR2") || args[0].ends_with(".dng") || args[0].ends_with(".jpg") {
"align".to_string()
} else {
args.remove(0)
};
let paths: Vec<PathBuf> = args.iter().map(PathBuf::from).collect();
// Previews, oriented, at proxy size.
let t = Instant::now();
let mut proxies: Vec<Gray> = Vec::new();
for p in &paths {
let bytes = std::fs::read(p).expect("read");
let preview = dr_decode::extract_preview(&bytes, dr_decode::PreviewSize::Full)
.expect("embedded preview");
let orientation =
dr_decode::orientation(&bytes[..bytes.len().min(dr_decode::HEADER_BYTES as usize)])
.unwrap_or(dr_types::Orientation::NORMAL);
let tag = match orientation.quarter_turns {
1 => 6,
2 => 3,
3 => 8,
_ => 1,
};
let gray = Gray::from_rgba8(
&preview.rgba,
preview.width as usize,
preview.height as usize,
)
.oriented(tag);
let (fitted, _) = gray.fitted(
dr_pano::xfeat::INPUT_LONG_EDGE,
dr_pano::xfeat::INPUT_LONG_EDGE,
);
println!(
"{:<14} preview {}×{} orientation {} → proxy {}×{}",
p.file_name().unwrap().to_string_lossy(),
preview.width,
preview.height,
tag,
fitted.width,
fitted.height
);
proxies.push(fitted);
}
println!("previews in {:?}", t.elapsed());
// Keypoints.
let t = Instant::now();
let mut detector = XFeat::embedded().expect("model");
let features: Vec<_> = proxies
.iter()
.map(|g| detector.detect(g).expect("detect"))
.collect();
for (i, f) in features.iter().enumerate() {
println!("frame {i}: {} keypoints", f.len());
}
println!(
"detection in {:?} ({:?} per frame)",
t.elapsed(),
t.elapsed() / proxies.len() as u32
);
// Alignment.
let t = Instant::now();
let opts = AlignOptions::default();
let alignment = align(&features, &opts).expect("align");
println!("alignment in {:?}", t.elapsed());
println!(
"focal {:.1} px, long edge {} px ({:.1} mm on full frame), rms {:.3} px",
alignment.focal,
proxies[0].width.max(proxies[0].height),
alignment.focal * 36.0 / proxies[0].width.max(proxies[0].height) as f64,
alignment.rms_px
);
for l in &alignment.links {
println!(
" link {}–{}: {} inliers of {} matches",
l.i, l.j, l.inliers, l.matches
);
}
for (k, why) in &alignment.unaligned {
println!(" UNALIGNED frame {k}: {why}");
}
let root = alignment
.rotations
.iter()
.position(|r| *r == Some(dr_pano::linalg::Mat3::IDENTITY))
.unwrap_or(0);
for (k, r) in alignment.rotations.iter().enumerate() {
if let Some(r) = r {
// Yaw about y, pitch about x, roll about z, from the matrix's
// columns — enough to read a sweep by eye.
let yaw = r.0[0][2].atan2(r.0[2][2]).to_degrees();
let pitch = (-r.0[1][2]).asin().to_degrees();
let roll = r.0[1][0].atan2(r.0[1][1]).to_degrees();
println!(
" frame {k}: yaw {yaw:7.2}° pitch {pitch:6.2}° roll {roll:6.2}°{}",
if k == root { " (reference)" } else { "" }
);
}
}
if !alignment.is_complete() {
eprintln!("not drawing: the set is not fully aligned");
std::process::exit(1);
}
// Draw: a cylinder, averaged where frames overlap.
let t = Instant::now();
let cameras: Cameras = alignment.cameras();
let (fw, fh) = (proxies[0].width as f64, proxies[0].height as f64);
let scale = alignment.focal;
let bounds = dr_pano::projection::bounds(Projection::Cylindrical, scale, &cameras, (fw, fh))
.expect("bounds");
// Fit to 3000 px wide.
let out_w = 3000usize;
let px = bounds.width() / out_w as f64;
let out_h = (bounds.height() / px).ceil() as usize;
let mut sum = vec![0.0f32; out_w * out_h];
let mut count = vec![0u16; out_w * out_h];
for oy in 0..out_h {
for ox in 0..out_w {
let u = bounds.min_u + (ox as f64 + 0.5) * px;
let v = bounds.min_v + (oy as f64 + 0.5) * px;
let d = Projection::Cylindrical.to_direction(scale, u, v);
for (k, g) in proxies.iter().enumerate() {
let Some((x, y)) = cameras.project(k, d) else {
continue;
};
let (x, y) = (x + g.width as f64 / 2.0, y + g.height as f64 / 2.0);
if x < 0.0 || y < 0.0 || x >= g.width as f64 - 1.0 || y >= g.height as f64 - 1.0 {
continue;
}
let (x0, y0) = (x as usize, y as usize);
let (tx, ty) = ((x - x0 as f64) as f32, (y - y0 as f64) as f32);
let p = |xx: usize, yy: usize| g.data[yy * g.width + xx];
let val = (p(x0, y0) * (1.0 - tx) + p(x0 + 1, y0) * tx) * (1.0 - ty)
+ (p(x0, y0 + 1) * (1.0 - tx) + p(x0 + 1, y0 + 1) * tx) * ty;
sum[oy * out_w + ox] += val;
count[oy * out_w + ox] += 1;
}
}
}
let mut ppm = format!("P5\n{out_w} {out_h}\n255\n").into_bytes();
ppm.extend(sum.iter().zip(&count).map(|(s, c)| {
if *c == 0 {
0u8
} else {
((s / f32::from(*c)).clamp(0.0, 1.0) * 255.0) as u8
}
}));
let out = format!("{prefix}-cyl.pgm");
std::fs::write(&out, ppm).expect("write");
println!("wrote {out} ({out_w}×{out_h}) in {:?}", t.elapsed());
}
+459
View File
@@ -0,0 +1,459 @@
//! TRACES: FR-MRG-1 | FR-MRG-5
//! From features to cameras: the alignment of a whole set.
//!
//! 1. Match every pair of frames (`matching`).
//! 2. For each pair with enough matches, a robust homography
//! (`homography::ransac_homography`); a pair is a *link* when its inliers
//! pass Brown & Lowe's test, `n_inliers > 8 + 0.3 · n_matches`, which
//! is what separates a real overlap from a coincidence of descriptors.
//! 3. The focal length: the median of what the links' homographies imply,
//! or the caller's hint if none of them implies anything.
//! 4. A spanning tree over the links, strongest first, from the
//! best-connected frame; rotations chained along it.
//! 5. Bundle adjustment over every link's inliers (`bundle`).
//!
//! What it refuses to do is guess. A frame the tree does not reach is
//! reported by index with the reason (FR-MRG-5) and left out of the
//! cameras; the caller decides whether a set with a hole is worth
//! stitching, and the requirement says it is not.
use crate::bundle::{self, AdjustOptions, Cameras, Observation};
use crate::features::Features;
use crate::homography::{self, RobustHomography};
use crate::linalg::Mat3;
use crate::matching::{match_features, Match};
use crate::PanoError;
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct AlignOptions {
/// Descriptor similarity floor for a match (`matching`).
pub min_similarity: f32,
/// RANSAC agreement distance, in pixels of the features' image.
pub ransac_px: f64,
pub ransac_iterations: usize,
/// A pair needs at least this many inliers to be a link, on top of
/// Brown & Lowe's ratio test.
pub min_inliers: usize,
/// Focal length in pixels of the features' image, if the caller knows
/// it (EXIF and a sensor width). Used only when the homographies do not
/// determine one.
pub focal_hint: Option<f64>,
pub adjust: AdjustOptions,
/// For RANSAC's sampling: the same seed gives the same alignment
/// (NFR-MRG-2).
pub seed: u64,
}
impl Default for AlignOptions {
fn default() -> Self {
AlignOptions {
min_similarity: 0.82,
ransac_px: 3.0,
ransac_iterations: 1000,
min_inliers: 12,
focal_hint: None,
adjust: AdjustOptions::default(),
seed: 0x5eed,
}
}
}
/// An overlap the alignment trusts.
#[derive(Debug, Clone, PartialEq)]
pub struct Link {
pub i: usize,
pub j: usize,
pub matches: usize,
pub inliers: usize,
/// Maps centred points of `i` to centred points of `j`.
pub h: Mat3,
}
/// Why a frame is not in the alignment.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum Unaligned {
/// Not enough matches with any other frame to try a geometry.
NoMatches,
/// Matches existed but none survived RANSAC as a real overlap.
NoOverlap,
/// Overlaps existed but only with frames that are themselves unaligned.
Disconnected,
}
impl std::fmt::Display for Unaligned {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str(match self {
Unaligned::NoMatches => "too few matching features with any other frame",
Unaligned::NoOverlap => "no consistent overlap with any other frame",
Unaligned::Disconnected => "overlaps only with frames that could not be aligned",
})
}
}
/// The result: cameras for the aligned frames, and the rest named.
#[derive(Debug, Clone, PartialEq)]
pub struct Alignment {
/// One rotation per input frame, camera to world, for aligned frames;
/// `None` for the unaligned. The reference frame is the best-connected
/// one and has the identity.
pub rotations: Vec<Option<Mat3>>,
/// Focal length in pixels of the features' image.
pub focal: f64,
pub links: Vec<Link>,
pub unaligned: Vec<(usize, Unaligned)>,
/// Bundle adjustment's RMS reprojection error, in pixels.
pub rms_px: f64,
}
impl Alignment {
pub fn is_complete(&self) -> bool {
self.unaligned.is_empty()
}
/// The cameras of the aligned frames, indexed as the input — a frame
/// that is not aligned is given the identity, so this is only useful
/// when [`Self::is_complete`].
pub fn cameras(&self) -> Cameras {
Cameras {
rotations: self
.rotations
.iter()
.map(|r| r.unwrap_or(Mat3::IDENTITY))
.collect(),
focal: self.focal,
}
}
}
/// Align a set of frames from their features.
///
/// Every `Features` must be in its own frame's pixel coordinates with the
/// image size filled in; points are centred on the image centre here. The
/// frames must all come from the same lens at the same focal length, which
/// is the panorama assumption and not checked — the caller has the EXIF.
pub fn align(frames: &[Features], opts: &AlignOptions) -> Result<Alignment, PanoError> {
let n = frames.len();
if n < 2 {
return Err(PanoError::Input(
"a panorama needs at least two frames".into(),
));
}
let centre = |k: usize, i: usize| -> (f64, f64) {
let kp = frames[k].keypoints[i];
(
f64::from(kp.x) - frames[k].width as f64 / 2.0,
f64::from(kp.y) - frames[k].height as f64 / 2.0,
)
};
// Scale for the DLT's conditioning: points of order one.
let scale = 1.0
/ frames
.iter()
.map(|f| f.width.max(f.height) as f64)
.fold(1.0, f64::max);
// 1 + 2: every pair.
let mut links = Vec::new();
let mut observations: Vec<Observation> = Vec::new();
let mut matched_any = vec![false; n];
let t_match = std::time::Instant::now();
for i in 0..n {
for j in i + 1..n {
let matches: Vec<Match> = match_features(&frames[i], &frames[j], opts.min_similarity);
log::debug!("pair {i}-{j}: {} matches", matches.len());
if matches.len() < 4 {
continue;
}
matched_any[i] = true;
matched_any[j] = true;
let pairs: Vec<((f64, f64), (f64, f64))> = matches
.iter()
.map(|m| {
let (a, b) = (centre(i, m.a), centre(j, m.b));
((a.0 * scale, a.1 * scale), (b.0 * scale, b.1 * scale))
})
.collect();
let Some(RobustHomography { h, inliers }) = homography::ransac_homography(
&pairs,
opts.ransac_px * scale,
opts.ransac_iterations,
opts.seed ^ ((i as u64) << 32 | j as u64),
) else {
continue;
};
let needed = (8.0 + 0.3 * matches.len() as f64).ceil() as usize;
log::debug!("pair {i}-{j}: {} inliers, {needed} needed", inliers.len());
if inliers.len() <= needed || inliers.len() < opts.min_inliers {
continue;
}
// Back to pixels: H_px = S⁻¹ H S.
let m = h.0;
let h_px = Mat3([
[m[0][0], m[0][1], m[0][2] / scale],
[m[1][0], m[1][1], m[1][2] / scale],
[m[2][0] * scale, m[2][1] * scale, m[2][2]],
]);
for &k in &inliers {
let (a, b) = pairs[k];
observations.push(Observation {
i,
j,
pi: (a.0 / scale, a.1 / scale),
pj: (b.0 / scale, b.1 / scale),
});
}
links.push(Link {
i,
j,
matches: matches.len(),
inliers: inliers.len(),
h: h_px,
});
}
}
log::debug!("matching and pairwise geometry in {:?}", t_match.elapsed());
// 3: the focal length.
let mut estimates: Vec<f64> = links
.iter()
.filter_map(|l| homography::focal_from_homography(&l.h))
.filter(|f| f.is_finite() && *f > 0.0)
.collect();
let longest = frames
.iter()
.map(|f| f.width.max(f.height) as f64)
.fold(0.0, f64::max);
let focal = if !estimates.is_empty() {
estimates.sort_by(f64::total_cmp);
let median = estimates[estimates.len() / 2];
// A homography of a nearly pure pan can imply almost anything;
// clamp to the range a real lens on this sensor can reach.
median.clamp(0.3 * longest, 6.0 * longest)
} else if let Some(hint) = opts.focal_hint {
hint
} else {
// No overlap said anything and nobody told us: a normal lens.
longest
};
// 4: spanning tree, strongest link first, from the best-connected frame.
let mut rotations: Vec<Option<Mat3>> = vec![None; n];
let mut unaligned = Vec::new();
if links.is_empty() {
for (k, &matched) in matched_any.iter().enumerate() {
unaligned.push((
k,
if matched {
Unaligned::NoOverlap
} else {
Unaligned::NoMatches
},
));
}
return Ok(Alignment {
rotations,
focal,
links,
unaligned,
rms_px: 0.0,
});
}
let mut degree = vec![0usize; n];
for l in &links {
degree[l.i] += l.inliers;
degree[l.j] += l.inliers;
}
let root = (0..n).max_by_key(|&k| degree[k]).unwrap_or(0);
rotations[root] = Some(Mat3::IDENTITY);
loop {
// The strongest link from an aligned frame to an unaligned one.
let best = links
.iter()
.filter(|l| rotations[l.i].is_some() != rotations[l.j].is_some())
.max_by_key(|l| l.inliers);
let Some(l) = best else { break };
let r_ij = homography::rotation_from_homography(&l.h, focal);
// H_ij takes points of i to j, so bearings b_j = R_ij b_i, and with
// world = R_i · cam_i: R_j = R_i · R_ijᵀ.
if let Some(ri) = rotations[l.i] {
rotations[l.j] = Some((ri * r_ij.transpose()).orthonormalised());
} else if let Some(rj) = rotations[l.j] {
rotations[l.i] = Some((rj * r_ij).orthonormalised());
}
}
for k in 0..n {
if rotations[k].is_none() {
let reason = if !matched_any[k] {
Unaligned::NoMatches
} else if links.iter().any(|l| l.i == k || l.j == k) {
Unaligned::Disconnected
} else {
Unaligned::NoOverlap
};
unaligned.push((k, reason));
}
}
// 5: adjust the aligned frames together. The reference frame must be
// index 0 of the adjustment (it holds frame 0 fixed), so the aligned
// frames are renumbered with the root first.
let aligned: Vec<usize> = std::iter::once(root)
.chain((0..n).filter(|&k| k != root && rotations[k].is_some()))
.collect();
let index_of = |k: usize| aligned.iter().position(|&a| a == k);
let start = Cameras {
rotations: aligned.iter().map(|&k| rotations[k].unwrap()).collect(),
focal,
};
let obs: Vec<Observation> = observations
.iter()
.filter_map(|o| {
Some(Observation {
i: index_of(o.i)?,
j: index_of(o.j)?,
pi: o.pi,
pj: o.pj,
})
})
.collect();
let t_adjust = std::time::Instant::now();
let adjusted = bundle::adjust(start, &obs, &opts.adjust)?;
log::debug!(
"bundle adjustment: {} observations, {} iterations in {:?}",
obs.len(),
adjusted.iterations,
t_adjust.elapsed()
);
for (slot, &k) in aligned.iter().enumerate() {
rotations[k] = Some(adjusted.cameras.rotations[slot]);
}
Ok(Alignment {
rotations,
focal: adjusted.cameras.focal,
links,
unaligned,
rms_px: adjusted.rms_px,
})
}
#[cfg(test)]
mod tests {
use super::*;
use crate::features::{Keypoint, DESCRIPTOR_LEN};
use crate::linalg::Vec3;
/// Frames of a synthetic sweep: world directions with random unit
/// descriptors, each frame seeing the ones in its field of view.
fn synthetic_sweep(
n: usize,
step: f64,
f: f64,
w: usize,
h: usize,
) -> (Vec<Features>, Cameras) {
let mut seed = 777u64;
let mut rnd = || {
seed = seed
.wrapping_mul(6364136223846793005)
.wrapping_add(1442695040888963407);
((seed >> 33) as f64 / (1u64 << 31) as f64) - 0.5
};
let rotations: Vec<Mat3> = (0..n)
.map(|k| {
Mat3::rotation(Vec3::new(0.0, 1.0, 0.0), step * k as f64)
* Mat3::rotation(Vec3::new(1.0, 0.0, 0.0), 0.02 * ((k % 3) as f64 - 1.0))
})
.collect();
let truth = Cameras {
rotations,
focal: f,
};
let total = step * (n as f64 - 1.0);
let mut frames: Vec<Features> = (0..n)
.map(|_| Features {
keypoints: Vec::new(),
descriptors: Vec::new(),
width: w,
height: h,
})
.collect();
for _ in 0..600 * n {
let yaw = rnd() * (total + 0.8) + total / 2.0;
let pitch = rnd() * 0.5;
let d = Vec3::new(
yaw.sin() * pitch.cos(),
pitch.sin(),
yaw.cos() * pitch.cos(),
);
let desc: Vec<f32> = (0..DESCRIPTOR_LEN).map(|_| rnd() as f32).collect();
let norm = desc.iter().map(|v| v * v).sum::<f32>().sqrt();
let desc: Vec<f32> = desc.iter().map(|v| v / norm).collect();
for (k, frame) in frames.iter_mut().enumerate() {
if let Some(p) = truth.project(k, d) {
let (x, y) = (p.0 + w as f64 / 2.0, p.1 + h as f64 / 2.0);
if x >= 0.0 && x < w as f64 && y >= 0.0 && y < h as f64 {
frame.keypoints.push(Keypoint {
x: (x + rnd() * 0.6) as f32,
y: (y + rnd() * 0.6) as f32,
score: 1.0,
});
frame.descriptors.extend_from_slice(&desc);
}
}
}
}
(frames, truth)
}
fn angle_between(a: Mat3, b: Mat3) -> f64 {
(a.transpose() * b).log().norm()
}
#[test]
fn a_synthetic_sweep_is_aligned_to_its_truth() {
let (frames, truth) = synthetic_sweep(6, 0.3, 1400.0, 1024, 768);
let out = align(&frames, &AlignOptions::default()).expect("aligned");
assert!(out.is_complete(), "unaligned: {:?}", out.unaligned);
assert_eq!(out.links.len(), 5 + 4, "links: {}", out.links.len());
assert!((out.focal - 1400.0).abs() < 15.0, "focal {}", out.focal);
assert!(out.rms_px < 1.0, "rms {}", out.rms_px);
// Relative rotations match the truth's, whichever frame is the root.
let root = out
.rotations
.iter()
.position(|r| *r == Some(Mat3::IDENTITY))
.unwrap();
for k in 0..6 {
let rel_truth = truth.rotations[root].transpose() * truth.rotations[k];
let rel_out = out.rotations[k].unwrap();
let err = angle_between(rel_truth, rel_out);
assert!(err < 2e-3, "frame {k} off by {err} rad");
}
}
#[test]
fn a_frame_from_nowhere_is_named_not_guessed() {
let (mut frames, _) = synthetic_sweep(4, 0.3, 1400.0, 1024, 768);
// Frame 3 gets descriptors nobody else has.
for v in &mut frames[3].descriptors {
*v = -*v;
}
let out = align(&frames, &AlignOptions::default()).expect("aligned");
assert_eq!(out.unaligned.len(), 1);
assert_eq!(out.unaligned[0].0, 3);
assert!(out.rotations[3].is_none());
assert!(out.rotations[..3].iter().all(Option::is_some));
}
#[test]
fn one_frame_is_refused() {
let (frames, _) = synthetic_sweep(1, 0.3, 1400.0, 640, 480);
assert!(matches!(
align(&frames, &AlignOptions::default()),
Err(PanoError::Input(_))
));
}
}
+446
View File
@@ -0,0 +1,446 @@
//! Bundle adjustment: every rotation and the focal length, refined together.
//!
//! The pairwise homographies (`homography.rs`) each know about two frames.
//! Chained around a loop they disagree with themselves by the accumulated
//! error, and a twelve-frame sweep chained end to end drifts by a visible
//! amount. This solves for all the rotations at once, against every inlier
//! match of every pair, so the error is spread rather than accumulated —
//! Brown & Lowe's step 4, with the camera model reduced to what a panorama
//! needs: one rotation per frame and one focal length shared by all.
//!
//! Levenberg–Marquardt with a numerical Jacobian. Analytic derivatives of a
//! rotation's projection are not hard, but they are a second place the
//! model is written down, and the model is small: forty parameters, a few
//! thousand residuals, a Jacobian that costs forty residual evaluations.
//! The whole solve is milliseconds. Correctness over cleverness, and one
//! definition of the projection to keep right.
use crate::linalg::{DMat, Mat3, Vec3};
use crate::PanoError;
/// A point in one image, centred on the principal point, in pixels.
pub type Point = (f64, f64);
/// One inlier match between two frames.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Observation {
pub i: usize,
pub j: usize,
pub pi: Point,
pub pj: Point,
}
/// What the adjustment starts from and returns: a rotation per frame
/// (camera to world; frame 0 is the world) and the focal length in pixels.
#[derive(Debug, Clone, PartialEq)]
pub struct Cameras {
pub rotations: Vec<Mat3>,
pub focal: f64,
}
impl Cameras {
/// The unit direction, in world space, that pixel `p` of frame `i` looks
/// along.
pub fn bearing(&self, i: usize, p: Point) -> Vec3 {
self.rotations[i] * Vec3::new(p.0, p.1, self.focal).normalised()
}
/// Where world direction `d` lands in frame `j`, or `None` if it is
/// behind the camera.
pub fn project(&self, j: usize, d: Vec3) -> Option<Point> {
let c = self.rotations[j].transpose() * d;
if c.z() <= 1e-9 {
return None;
}
Some((self.focal * c.x() / c.z(), self.focal * c.y() / c.z()))
}
}
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct AdjustOptions {
pub max_iterations: usize,
/// Residuals beyond this many pixels are down-weighted (Huber), so a
/// mismatch RANSAC let through pulls with bounded force.
pub huber_px: f64,
/// Whether the focal length is a free parameter. Off, it is held at the
/// starting value — for a set whose rotations are all small, the focal
/// length is weakly observable and better taken from the homographies'
/// median than pulled about by noise.
pub refine_focal: bool,
}
impl Default for AdjustOptions {
fn default() -> Self {
AdjustOptions {
max_iterations: 50,
huber_px: 3.0,
refine_focal: true,
}
}
}
/// The adjusted cameras and the fit.
#[derive(Debug, Clone, PartialEq)]
pub struct Adjusted {
pub cameras: Cameras,
/// Root-mean-square reprojection error over all observations, in pixels
/// (unweighted, so an outlier RANSAC missed shows here rather than
/// hiding under its Huber weight).
pub rms_px: f64,
pub iterations: usize,
}
/// Refine `start` against `observations`.
///
/// Frame 0's rotation is held fixed: the world frame is arbitrary and
/// fixing one camera removes the freedom. Every other frame must appear in
/// at least one observation or its rotation is undetermined and the normal
/// equations are singular — the caller (`align`) guarantees it by only
/// adjusting frames a spanning tree reached.
pub fn adjust(
start: Cameras,
observations: &[Observation],
opts: &AdjustOptions,
) -> Result<Adjusted, PanoError> {
let n_frames = start.rotations.len();
if n_frames < 2 || observations.is_empty() {
let rms = rms(&start, observations);
return Ok(Adjusted {
cameras: start,
rms_px: rms,
iterations: 0,
});
}
// Every adjustable frame must be constrained by something, or its
// block of the normal equations is zero and the solve is meaningless —
// checked here, by name, rather than left to surface as a step that
// fails to lower the cost.
let mut seen = vec![false; n_frames];
for o in observations {
seen[o.i] = true;
seen[o.j] = true;
}
if let Some(k) = (1..n_frames).find(|&k| !seen[k]) {
return Err(PanoError::Geometry(format!(
"frame {k} has no observations constraining it"
)));
}
let n_rot = 3 * (n_frames - 1);
let n_params = n_rot + usize::from(opts.refine_focal);
let n_res = 2 * observations.len();
// Parameters are *increments* on the current cameras, re-applied each
// accepted step: rotation k ← exp(δ_k) · rotation k, focal ← f · exp(δ_f).
// Composing on the left keeps the increment in world space, where a
// small rotation means the same thing for every frame.
let apply = |base: &Cameras, x: &[f64]| -> Cameras {
let mut rotations = base.rotations.clone();
for k in 1..n_frames {
let w = Vec3::new(x[3 * (k - 1)], x[3 * (k - 1) + 1], x[3 * (k - 1) + 2]);
rotations[k] = (Mat3::exp(w) * base.rotations[k]).orthonormalised();
}
let focal = if opts.refine_focal {
base.focal * x[n_rot].exp()
} else {
base.focal
};
Cameras { rotations, focal }
};
let residuals = |c: &Cameras, out: &mut Vec<f64>| {
out.clear();
for o in observations {
let d = c.bearing(o.i, o.pi);
match c.project(o.j, d) {
Some((x, y)) => {
out.push(x - o.pj.0);
out.push(y - o.pj.1);
}
None => {
// Behind the camera: as wrong as a residual can be
// without being infinite. The Huber weight caps its pull.
out.push(1e4);
out.push(1e4);
}
}
}
};
let weights = |r: &[f64], out: &mut Vec<f64>| {
out.clear();
for pair in r.chunks_exact(2) {
let m = (pair[0] * pair[0] + pair[1] * pair[1]).sqrt();
let w = if m > opts.huber_px {
opts.huber_px / m
} else {
1.0
};
out.push(w);
out.push(w);
}
};
// The robust cost itself, not the weighted sum of squares: the weights
// above are the IRLS linearisation for one step, and comparing two
// steps by sums taken under different weights would accept the wrong
// ones. Huber: quadratic within the threshold, linear beyond it.
let cost = |r: &[f64]| -> f64 {
r.chunks_exact(2)
.map(|pair| {
let m = (pair[0] * pair[0] + pair[1] * pair[1]).sqrt();
if m <= opts.huber_px {
m * m
} else {
2.0 * opts.huber_px * m - opts.huber_px * opts.huber_px
}
})
.sum()
};
let mut cameras = start;
let mut r = Vec::with_capacity(n_res);
let mut w = Vec::with_capacity(n_res);
residuals(&cameras, &mut r);
weights(&r, &mut w);
let mut current = cost(&r);
let mut lambda = 1e-3;
let mut jac = vec![0.0f64; n_res * n_params];
let mut r_plus = Vec::with_capacity(n_res);
let zero = vec![0.0f64; n_params];
let mut iterations = 0;
for _ in 0..opts.max_iterations {
iterations += 1;
// Numerical Jacobian about the current cameras (x = 0).
const H: f64 = 1e-6;
for p in 0..n_params {
let mut x = zero.clone();
x[p] = H;
let c_plus = apply(&cameras, &x);
residuals(&c_plus, &mut r_plus);
for (k, (rp, r0)) in r_plus.iter().zip(&r).enumerate() {
jac[k * n_params + p] = (rp - r0) / H;
}
}
// Normal equations, weighted: (JᵀWJ + λ·diag) δ = −JᵀWr.
let mut a = DMat::zeros(n_params);
let mut b = vec![0.0f64; n_params];
for k in 0..n_res {
let row = &jac[k * n_params..(k + 1) * n_params];
let wk = w[k];
for p in 0..n_params {
b[p] -= wk * row[p] * r[k];
for q in 0..n_params {
a[(p, q)] += wk * row[p] * row[q];
}
}
}
// Try steps with increasing damping until one lowers the cost.
let mut accepted = false;
for _ in 0..10 {
let mut damped = a.clone();
for p in 0..n_params {
let d = a[(p, p)];
damped[(p, p)] = d + lambda * d.max(1e-9);
}
let Some(delta) = damped.solve_spd(&b) else {
return Err(PanoError::Geometry(
"the adjustment's normal equations are singular: a frame has no \
observations constraining it"
.into(),
));
};
let candidate = apply(&cameras, &delta);
residuals(&candidate, &mut r_plus);
let c_new = cost(&r_plus);
if c_new < current {
let improvement = (current - c_new) / current.max(1e-12);
let step: f64 = delta.iter().map(|d| d * d).sum::<f64>().sqrt();
cameras = candidate;
std::mem::swap(&mut r, &mut r_plus);
weights(&r, &mut w);
current = c_new;
lambda = (lambda / 3.0).max(1e-9);
accepted = true;
// Converged when a *lightly damped* step no longer helps. A
// heavily damped step is small by construction and would
// pass an improvement test long before the minimum.
if step < 1e-10 || (improvement < 1e-8 && lambda < 1e-2) {
return Ok(Adjusted {
rms_px: rms(&cameras, observations),
cameras,
iterations,
});
}
break;
}
lambda *= 5.0;
}
if !accepted {
break;
}
}
Ok(Adjusted {
rms_px: rms(&cameras, observations),
cameras,
iterations,
})
}
/// Unweighted RMS reprojection error in pixels.
pub fn rms(c: &Cameras, observations: &[Observation]) -> f64 {
if observations.is_empty() {
return 0.0;
}
let sum: f64 = observations
.iter()
.map(|o| match c.project(o.j, c.bearing(o.i, o.pi)) {
Some((x, y)) => (x - o.pj.0).powi(2) + (y - o.pj.1).powi(2),
None => 1e8,
})
.sum();
(sum / observations.len() as f64).sqrt()
}
#[cfg(test)]
mod tests {
use super::*;
/// A synthetic sweep: `n` cameras panned by `step` radians each with a
/// little pitch and roll, `f` pixels, and matches between neighbours
/// from a cloud of world directions.
fn sweep(n: usize, step: f64, f: f64, noise_px: f64) -> (Cameras, Vec<Observation>) {
let mut rotations = Vec::new();
for k in 0..n {
let yaw = step * k as f64;
let pitch = 0.01 * ((k * 7) % 3) as f64;
let roll = 0.005 * ((k * 5) % 4) as f64;
let r = Mat3::rotation(Vec3::new(0.0, 1.0, 0.0), yaw)
* Mat3::rotation(Vec3::new(1.0, 0.0, 0.0), pitch)
* Mat3::rotation(Vec3::new(0.0, 0.0, 1.0), roll);
rotations.push(r);
}
let truth = Cameras {
rotations,
focal: f,
};
// World directions: a fan across the whole sweep.
let mut obs = Vec::new();
let mut seed = 12345u64;
let mut rnd = || {
seed = seed
.wrapping_mul(6364136223846793005)
.wrapping_add(1442695040888963407);
((seed >> 33) as f64 / (1u64 << 31) as f64) - 0.5
};
let total = step * (n as f64 - 1.0);
for _ in 0..400 * n {
let yaw = rnd() * (total + 0.8) + total / 2.0;
let pitch = rnd() * 0.5;
let d = Vec3::new(
yaw.sin() * pitch.cos(),
pitch.sin(),
yaw.cos() * pitch.cos(),
)
.normalised();
// Visible in which frames? Within ±0.35 f of centre.
let mut seen: Vec<(usize, Point)> = Vec::new();
for k in 0..n {
if let Some(p) = truth.project(k, d) {
if p.0.abs() < 0.35 * f && p.1.abs() < 0.25 * f {
seen.push((k, (p.0 + rnd() * noise_px, p.1 + rnd() * noise_px)));
}
}
}
for a in 0..seen.len() {
for b in a + 1..seen.len() {
obs.push(Observation {
i: seen[a].0,
j: seen[b].0,
pi: seen[a].1,
pj: seen[b].1,
});
}
}
}
(truth, obs)
}
fn angle_between(a: Mat3, b: Mat3) -> f64 {
(a.transpose() * b).log().norm()
}
#[test]
fn a_perturbed_start_converges_back_to_the_truth() {
let (truth, obs) = sweep(6, 0.3, 1400.0, 0.0);
assert!(obs.len() > 500);
// Perturb every rotation but the first by ~1°, and the focal by 5%.
let mut start = truth.clone();
for k in 1..6 {
let w = Vec3::new(0.01, -0.015, 0.008) * (k as f64 / 3.0);
start.rotations[k] = Mat3::exp(w) * start.rotations[k];
}
start.focal *= 1.05;
let before = rms(&start, &obs);
let out = adjust(start, &obs, &AdjustOptions::default()).expect("solvable");
assert!(out.rms_px < 1e-3, "rms {} (was {before})", out.rms_px);
assert!(
(out.cameras.focal - 1400.0).abs() < 0.5,
"focal {}",
out.cameras.focal
);
for k in 0..6 {
let err = angle_between(out.cameras.rotations[k], truth.rotations[k]);
assert!(err < 1e-5, "frame {k} off by {err} rad");
}
}
#[test]
fn noise_is_averaged_rather_than_accumulated() {
let (truth, obs) = sweep(8, 0.25, 1400.0, 1.0);
let mut start = truth.clone();
for k in 1..8 {
start.rotations[k] =
Mat3::exp(Vec3::new(0.0, 0.004 * k as f64, 0.0)) * start.rotations[k];
}
let out = adjust(start, &obs, &AdjustOptions::default()).expect("solvable");
// ±0.5 px of uniform noise on every coordinate has an RMS of 0.41 px
// per axis, so the fit's RMS over both axes should sit near 0.58 and
// cannot be much below it.
assert!(out.rms_px < 0.7, "rms {}", out.rms_px);
// The focal length and the sweep are nearly degenerate for a
// single row: only the perspective inside each overlap pins the
// focal, and a pixel of noise is worth about a tenth of a percent of
// it. What that error does is scale every yaw by the same factor —
// a uniform stretch of the panorama, invisible in the result — so the
// absolute rotation error grows linearly along the sweep and is not
// the measure of the solve. The residual after removing that stretch
// is.
let f_ratio = out.cameras.focal / 1400.0;
assert!((f_ratio - 1.0).abs() < 5e-3, "focal {}", out.cameras.focal);
for k in 0..8 {
let yaw_k = 0.25 * k as f64;
let expected_stretch = (f_ratio - 1.0).abs() * yaw_k;
let err = angle_between(out.cameras.rotations[k], truth.rotations[k]);
assert!(
err < expected_stretch + 1.5e-4,
"frame {k} off by {err} rad, {expected_stretch} of it the focal's"
);
}
}
#[test]
fn a_frame_without_observations_is_refused() {
let (truth, mut obs) = sweep(4, 0.3, 1400.0, 0.0);
obs.retain(|o| o.i != 3 && o.j != 3);
let err = adjust(truth, &obs, &AdjustOptions::default()).unwrap_err();
assert!(matches!(err, PanoError::Geometry(_)));
}
}
+336
View File
@@ -0,0 +1,336 @@
//! Keypoints with descriptors, and the decoder that reads them out of
//! XFeat's dense maps.
//!
//! The network (S15.2) produces three maps at an eighth of the input
//! resolution and stops; everything from there to a list of keypoints is
//! this file, in plain Rust, for the reason `dr-segment` decodes yolo26's
//! heads itself: the post-processing is cheap, shape-dependent and exactly
//! the kind of graph tract parses badly. It is a port of the reference
//! `XFeat.detectAndCompute`, step for step, so that a keypoint here is the
//! keypoint the paper's numbers were measured on.
/// One detected point, in the pixel coordinates of the image it was
/// detected in, with the detector's confidence.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Keypoint {
pub x: f32,
pub y: f32,
/// The reliability the detector assigned; higher is better, and the
/// scale is the detector's own — comparable within one model only.
pub score: f32,
}
/// The keypoints of one image and their descriptors.
#[derive(Debug, Clone, PartialEq)]
pub struct Features {
pub keypoints: Vec<Keypoint>,
/// `keypoints.len() × DESCRIPTOR_LEN`, each row L2-normalised, so that a
/// dot product between two rows is their cosine similarity.
pub descriptors: Vec<f32>,
/// The image the coordinates are in.
pub width: usize,
pub height: usize,
}
/// The length of one descriptor. XFeat's is 64; the matcher does not care
/// what the number is, only that both sides agree.
pub const DESCRIPTOR_LEN: usize = 64;
impl Features {
pub fn len(&self) -> usize {
self.keypoints.len()
}
pub fn is_empty(&self) -> bool {
self.keypoints.is_empty()
}
pub fn descriptor(&self, i: usize) -> &[f32] {
&self.descriptors[i * DESCRIPTOR_LEN..(i + 1) * DESCRIPTOR_LEN]
}
}
/// XFeat's three output maps, as the network hands them back.
///
/// All three are `channels × height × width` at an eighth of the input, in
/// the NCHW order the ONNX export declares (`feats [1, 64, H/8, W/8]`,
/// `keypoints [1, 65, H/8, W/8]`, `heatmap [1, 1, H/8, W/8]`).
pub struct XFeatMaps<'a> {
/// 64 channels: the dense descriptor field.
pub feats: &'a [f32],
/// 65 channels: for each 8×8 cell, a logit per position plus one for
/// "no keypoint here".
pub keypoints: &'a [f32],
/// 1 channel: reliability.
pub heatmap: &'a [f32],
/// The maps' width and height (the input's, divided by eight).
pub width: usize,
pub height: usize,
}
/// How the decoder picks keypoints.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct DecodeOptions {
/// Keep at most this many, by score. The reference default is 4096.
pub top_k: usize,
/// A cell position's softmax probability must exceed this to be a
/// keypoint at all. The reference default is 0.05.
pub threshold: f32,
/// Ignore keypoints within this many pixels of the map's edge. A frame
/// padded into the detector's fixed input (`Gray::padded`) has a hard
/// edge where the padding starts, and the detector fires on it.
pub border: usize,
}
impl Default for DecodeOptions {
fn default() -> Self {
DecodeOptions {
top_k: 4096,
threshold: 0.05,
border: 4,
}
}
}
/// Decode keypoints and descriptors from the network's maps.
///
/// The reference, step for step:
/// 1. softmax over the 65 logits of each cell, keep the 64 positions;
/// 2. pixel-shuffle those into a full-resolution keypoint heatmap — channel
/// `c` of cell `(cx, cy)` is pixel `(cx·8 + c%8, cy·8 + c/8)`;
/// 3. 5×5 non-maximum suppression over that heatmap, above `threshold`;
/// 4. score each survivor by its heatmap value times the reliability map
/// sampled bilinearly at its position;
/// 5. keep the `top_k` by score;
/// 6. sample the descriptor field bilinearly at each and L2-normalise.
///
/// Bilinear where the reference samples the descriptor field bicubically:
/// a quarter-pixel's difference in a field that is smooth by construction,
/// and one interpolator rather than two to keep correct.
pub fn decode_xfeat(maps: &XFeatMaps<'_>, opts: &DecodeOptions) -> Features {
let (w8, h8) = (maps.width, maps.height);
let (w, h) = (w8 * 8, h8 * 8);
let cells = w8 * h8;
debug_assert_eq!(maps.keypoints.len(), 65 * cells);
debug_assert_eq!(maps.feats.len(), DESCRIPTOR_LEN * cells);
debug_assert_eq!(maps.heatmap.len(), cells);
// 1 + 2: softmax per cell, scattered into the full-resolution heatmap.
let mut heat = vec![0.0f32; w * h];
for cy in 0..h8 {
for cx in 0..w8 {
let cell = cy * w8 + cx;
let logit = |c: usize| maps.keypoints[c * cells + cell];
let max = (0..65).map(logit).fold(f32::MIN, f32::max);
let mut sum = 0.0f32;
let mut exps = [0.0f32; 65];
for (c, e) in exps.iter_mut().enumerate() {
*e = (logit(c) - max).exp();
sum += *e;
}
for (c, e) in exps.iter().enumerate().take(64) {
let (dx, dy) = (c % 8, c / 8);
heat[(cy * 8 + dy) * w + cx * 8 + dx] = e / sum;
}
}
}
// 3: a pixel survives if it is the maximum of its 5×5 neighbourhood and
// above threshold. Ties go to every tied pixel, as the reference's
// `x == max_pool(x)` does.
let border = opts.border.max(2);
let mut survivors: Vec<(usize, usize, f32)> = Vec::new();
for y in border..h.saturating_sub(border) {
for x in border..w.saturating_sub(border) {
let v = heat[y * w + x];
if v <= opts.threshold {
continue;
}
let mut is_max = true;
'nb: for ny in y - 2..=y + 2 {
for nx in x - 2..=x + 2 {
if heat[ny * w + nx] > v {
is_max = false;
break 'nb;
}
}
}
if is_max {
survivors.push((x, y, v));
}
}
}
// 4: heatmap value × reliability, the latter sampled at the keypoint's
// position in map coordinates (`align_corners = False`: pixel `x` of the
// full image is `x / 8 - 0.5` in the map).
let sample = |field: &[f32], channels: usize, c: usize, x: f32, y: f32| -> f32 {
let fx = (x / 8.0 - 0.5).clamp(0.0, (w8 - 1) as f32);
let fy = (y / 8.0 - 0.5).clamp(0.0, (h8 - 1) as f32);
let x0 = fx as usize;
let y0 = fy as usize;
let x1 = (x0 + 1).min(w8 - 1);
let y1 = (y0 + 1).min(h8 - 1);
let tx = fx - x0 as f32;
let ty = fy - y0 as f32;
let at = |xx: usize, yy: usize| field[c * (w8 * h8) + yy * w8 + xx];
let _ = channels;
let top = at(x0, y0) * (1.0 - tx) + at(x1, y0) * tx;
let bot = at(x0, y1) * (1.0 - tx) + at(x1, y1) * tx;
top * (1.0 - ty) + bot * ty
};
let mut scored: Vec<(usize, usize, f32)> = survivors
.into_iter()
.map(|(x, y, v)| {
let r = sample(maps.heatmap, 1, 0, x as f32, y as f32);
(x, y, v * r)
})
.collect();
// 5: best first, then cut. `sort_unstable_by` on a total order of the
// score; NaN cannot occur — every input is a probability or a sigmoid.
scored.sort_unstable_by(|a, b| b.2.total_cmp(&a.2));
scored.truncate(opts.top_k);
// 6: descriptors.
let mut keypoints = Vec::with_capacity(scored.len());
let mut descriptors = Vec::with_capacity(scored.len() * DESCRIPTOR_LEN);
for (x, y, score) in scored {
let (xf, yf) = (x as f32, y as f32);
let start = descriptors.len();
for c in 0..DESCRIPTOR_LEN {
descriptors.push(sample(maps.feats, DESCRIPTOR_LEN, c, xf, yf));
}
let norm = descriptors[start..]
.iter()
.map(|v| v * v)
.sum::<f32>()
.sqrt()
.max(1e-12);
for v in &mut descriptors[start..] {
*v /= norm;
}
keypoints.push(Keypoint {
x: xf,
y: yf,
score,
});
}
Features {
keypoints,
descriptors,
width: w,
height: h,
}
}
#[cfg(test)]
mod tests {
use super::*;
/// Maps for a `w8 × h8` grid where every cell says "no keypoint" except
/// the listed ones, which put all their weight on one position.
fn maps(w8: usize, h8: usize, hot: &[(usize, usize, usize)]) -> (Vec<f32>, Vec<f32>, Vec<f32>) {
let cells = w8 * h8;
let mut kp = vec![0.0f32; 65 * cells];
// "None" strongly preferred everywhere.
for cell in 0..cells {
kp[64 * cells + cell] = 10.0;
}
for &(cx, cy, c) in hot {
let cell = cy * w8 + cx;
kp[64 * cells + cell] = 0.0;
kp[c * cells + cell] = 10.0;
}
let heat = vec![0.5f32; cells];
// Descriptors: channel c is constant c across the field, so any
// sampled descriptor is the same known vector.
let mut feats = vec![0.0f32; DESCRIPTOR_LEN * cells];
for c in 0..DESCRIPTOR_LEN {
for v in &mut feats[c * cells..(c + 1) * cells] {
*v = c as f32;
}
}
(feats, kp, heat)
}
#[test]
fn a_hot_cell_position_becomes_a_keypoint_at_the_right_pixel() {
// Cell (2, 1), channel 8*3 + 5 = 29 → pixel (2*8 + 5, 1*8 + 3).
let (f, k, h) = maps(8, 8, &[(2, 1, 29)]);
let out = decode_xfeat(
&XFeatMaps {
feats: &f,
keypoints: &k,
heatmap: &h,
width: 8,
height: 8,
},
&DecodeOptions::default(),
);
assert_eq!(out.len(), 1);
assert_eq!((out.keypoints[0].x, out.keypoints[0].y), (21.0, 11.0));
assert_eq!((out.width, out.height), (64, 64));
// Score is the softmax weight (~1) times the reliability (0.5).
assert!((out.keypoints[0].score - 0.5).abs() < 5e-3);
}
#[test]
fn descriptors_are_unit_length() {
let (f, k, h) = maps(8, 8, &[(3, 3, 0), (5, 5, 63)]);
let out = decode_xfeat(
&XFeatMaps {
feats: &f,
keypoints: &k,
heatmap: &h,
width: 8,
height: 8,
},
&DecodeOptions::default(),
);
assert_eq!(out.len(), 2);
for i in 0..2 {
let n: f32 = out.descriptor(i).iter().map(|v| v * v).sum();
assert!((n - 1.0).abs() < 1e-5);
}
}
#[test]
fn top_k_keeps_the_best() {
let (f, k, mut h) = maps(8, 8, &[(1, 1, 0), (3, 3, 0), (5, 5, 0)]);
// Make cell (3, 3) the most reliable.
h[3 * 8 + 3] = 0.9;
let out = decode_xfeat(
&XFeatMaps {
feats: &f,
keypoints: &k,
heatmap: &h,
width: 8,
height: 8,
},
&DecodeOptions {
top_k: 1,
..Default::default()
},
);
assert_eq!(out.len(), 1);
assert_eq!((out.keypoints[0].x, out.keypoints[0].y), (24.0, 24.0));
}
#[test]
fn the_border_is_excluded() {
let (f, k, h) = maps(8, 8, &[(0, 0, 0)]);
let out = decode_xfeat(
&XFeatMaps {
feats: &f,
keypoints: &k,
heatmap: &h,
width: 8,
height: 8,
},
&DecodeOptions::default(),
);
assert!(out.is_empty());
}
}
+793
View File
@@ -0,0 +1,793 @@
//! TRACES: FR-MRG-4
//! Filling a composite's uncovered border, tile by tile, with an inpainter.
//!
//! A merged panorama has a ragged border where no frame reached. FR-MRG-4
//! crops it by default; this fills it instead, when the photographer asks,
//! with pixels a model invents from the picture around them. Everything
//! here is the geometry of that — which tiles to run, what context to hand
//! the model, how to put its answers back — and none of it is the model:
//! that is the [`Inpainter`] trait, with MI-GAN behind it in `migan.rs`
//! and a fake in the tests.
//!
//! # Context across the edge
//!
//! An inpainting model is trained on holes *inside* pictures. A panorama's
//! border is a hole at the picture's *edge*: real content on one side,
//! nothing on the other, and a model given that invents a structure along
//! the open side — streaks of road in the sky, on the first try
//! (2026-09-19). So the known content is mirrored across the coverage
//! edge, column by column for the top and bottom bands and row by row for
//! the sides, into the hole and into a padding ring around the picture,
//! and the ring is presented as *known*. The model then interpolates
//! between real content and its mirror rather than extrapolating into
//! nothing. The ring is cut off at the end.
//!
//! # Structure from far away, texture from near
//!
//! One tiled pass at the working resolution was not enough: a 512-px tile
//! straddling the coverage edge sees a few hundred pixels of real content
//! on one side and invents the rest from that, two neighbouring tiles
//! invent differently, and the seams and the merge's own fringe leak into
//! the fill. [`fill_border`] therefore runs in two stages. A **coarse**
//! pass at a quarter of the size, where the whole border and hundreds of
//! pixels of real context sit inside a handful of tiles, decides the
//! structure — where the slope goes, where the sky stays sky. Then
//! **fine** passes regenerate the hole in bands from the real edge
//! outward: each band is the only unknown, with real content (or the band
//! before, freshly textured) on its near side and the coarse fill,
//! upsampled, on its far side — blurry, but the right structure — so the
//! model generates texture and a transition, never a large hole from
//! nothing.
//!
//! Tiles overlap by a third and are blended under a raised-cosine window,
//! so the seams between tiles do not show; the model's answer replaces
//! only the pixels that were unknown, and the picture itself is untouched.
use crate::PanoError;
/// A model that fills a square hole from its surroundings.
pub trait Inpainter {
/// The square tile it takes, in pixels.
fn tile(&self) -> usize;
/// Fill one tile. `rgb` is `tile × tile × 3`, row-major, 0..1, with the
/// unknown pixels' values meaningless; `known` is `tile × tile`. The
/// result is `tile × tile × 3`, 0..1, of which only the unknown pixels
/// are read.
fn fill(&mut self, rgb: &[f32], known: &[bool]) -> Result<Vec<f32>, PanoError>;
}
/// What a caller hears from [`fill_border`]: progress, for a page's bar,
/// and — for whoever is looking at why a fill went wrong — each stage's
/// picture as it lands. A plain `FnMut(usize, usize)` is an observer that
/// hears only the progress.
pub trait Observer {
/// `(done, total)` tiles, the total an estimate until the last band.
fn progress(&mut self, done: usize, total: usize);
/// A stage's result, `width × height × 3`: `coarse` (at the coarse
/// size), `band-N` after each fine band, `feathered` at the end.
fn stage(&mut self, _name: &str, _rgb: &[f32], _width: usize, _height: usize) {}
}
impl<F: FnMut(usize, usize)> Observer for F {
fn progress(&mut self, done: usize, total: usize) {
self(done, total)
}
}
/// How far the picture is extended with mirrored content before tiling.
/// Half a tile: enough that a hole at the edge sits well inside a tile.
pub const RING: usize = 256;
/// The fill's knobs, in pixels of the working image. The defaults are
/// what the fixture panorama looked best with on 2026-09-19; the merge
/// page exposes every one of them while the fill is experimental, so a
/// bad corner can be worked on from the picture rather than the code.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Params {
/// The coarse pass's reduction: 1 skips it.
pub coarse: usize,
/// The fine passes' band width.
pub band: usize,
/// How deep into the picture the mirrored context reaches. A plain
/// reflection of a deep hole pulls in whatever is that far from the
/// edge — a ridge, a peak — and the model, told that is what lies
/// beyond, paints it upside down. Folding the reflection within this
/// band keeps the ring looking like the edge it continues (sky beside
/// sky, grass beside grass) and nothing further away.
pub mirror_depth: usize,
/// How far inside the real edge the fill also regenerates, the two
/// blended by distance. A hard cut between real pixels and invented
/// ones is a line whatever the fill's quality; blended over this many
/// pixels it is not. Zero is the hard cut.
pub feather: usize,
/// The step between tiles, at most the tile; two thirds of it usual.
pub stride: usize,
}
impl Default for Params {
fn default() -> Self {
Params {
coarse: 4,
band: 96,
mirror_depth: 48,
feather: 24,
stride: 384,
}
}
}
/// Fill the unknown pixels of `rgb` (`width × height × 3`, 0..1) in place:
/// the coarse pass, then the fine bands, then the seam feathered over
/// `feather` pixels inside the real edge. Returns the tiles run.
///
/// `known` is `width × height`. `observer` hears the progress and, if it
/// cares, each stage.
pub fn fill_border(
rgb: &mut [f32],
width: usize,
height: usize,
known: &[bool],
model: &mut dyn Inpainter,
params: Params,
observer: &mut dyn Observer,
) -> Result<usize, PanoError> {
let Params {
coarse: q,
band,
mirror_depth,
feather,
stride,
} = params;
let q = q.max(1);
let band = band.max(8);
let mirror_depth = mirror_depth.max(1);
if width == 0 || height == 0 || rgb.len() != width * height * 3 || known.len() != width * height
{
return Err(PanoError::Input("fill: buffer sizes disagree".into()));
}
if known.iter().all(|&k| k) {
return Ok(0);
}
// The fill regenerates a margin inside the real edge too, and the
// result is blended with the real pixels across it at the end.
let real = rgb.to_vec();
let outer = known.to_vec();
let mut inner = known.to_vec();
erode(&mut inner, width, height, feather);
let known = &inner[..];
let mut done = 0usize;
// Coarse: a fraction of the size, unknown where any pixel of the cell was.
let (cw, ch) = ((width / q).max(1), (height / q).max(1));
let mut coarse = vec![0.0f32; cw * ch * 3];
let mut cknown = vec![true; cw * ch];
for y in 0..ch {
for x in 0..cw {
let mut sum = [0.0f32; 3];
let mut n = 0.0f32;
let mut all_known = true;
for dy in 0..q {
for dx in 0..q {
let (sx, sy) = ((x * q + dx).min(width - 1), (y * q + dy).min(height - 1));
let i = sy * width + sx;
all_known &= known[i];
for c in 0..3 {
sum[c] += rgb[i * 3 + c];
}
n += 1.0;
}
}
for c in 0..3 {
coarse[(y * cw + x) * 3 + c] = sum[c] / n;
}
cknown[y * cw + x] = all_known;
}
}
let estimate = |tiles: usize| tiles * 4;
done += fill_once(
&mut coarse,
cw,
ch,
&cknown,
model,
stride,
mirror_depth,
|n, t| observer.progress(n, estimate(t)),
)?;
observer.stage("coarse", &coarse, cw, ch);
// The hole starts as the coarse structure, upsampled.
for y in 0..height {
for x in 0..width {
let i = y * width + x;
if known[i] {
continue;
}
let fx = ((x as f32 + 0.5) / q as f32 - 0.5).clamp(0.0, (cw - 1) as f32);
let fy = ((y as f32 + 0.5) / q as f32 - 0.5).clamp(0.0, (ch - 1) as f32);
let (x0, y0) = (fx as usize, fy as usize);
let (x1, y1) = ((x0 + 1).min(cw - 1), (y0 + 1).min(ch - 1));
let (tx, ty) = (fx - x0 as f32, fy - y0 as f32);
for c in 0..3 {
let at = |xx: usize, yy: usize| coarse[(yy * cw + xx) * 3 + c];
rgb[i * 3 + c] = (at(x0, y0) * (1.0 - tx) + at(x1, y0) * tx) * (1.0 - ty)
+ (at(x0, y1) * (1.0 - tx) + at(x1, y1) * tx) * ty;
}
}
}
// Fine, in bands from the edge outward.
let dist = distance_to_known(known, width, height);
let mut band_known = vec![true; width * height];
let mut b = 0usize;
loop {
let lo = (b * band).saturating_sub(band / 2) as f32;
let hi = ((b + 1) * band) as f32;
let mut any = false;
for i in 0..width * height {
let in_band = !known[i] && dist[i] > lo && dist[i] <= hi;
band_known[i] = !in_band;
any |= in_band;
}
if !any {
break;
}
let before = done;
done += fill_once(
rgb,
width,
height,
&band_known,
model,
stride,
mirror_depth,
|n, t| observer.progress(before + n, before + estimate(t)),
)?;
observer.stage(&format!("band-{b}"), rgb, width, height);
b += 1;
}
// The seam: across the margin, real on the inside, invented on the
// outside, a smooth ramp between by distance from the true hole.
if feather > 0 {
let to_hole =
distance_to_known(&outer.iter().map(|k| !k).collect::<Vec<_>>(), width, height);
for i in 0..width * height {
if !outer[i] || known[i] {
continue;
}
// In the margin: outer says known, inner says not.
let t = (to_hole[i] / feather as f32).clamp(0.0, 1.0);
let t = t * t * (3.0 - 2.0 * t);
for c in 0..3 {
rgb[i * 3 + c] = rgb[i * 3 + c] * (1.0 - t) + real[i * 3 + c] * t;
}
}
}
observer.stage("feathered", rgb, width, height);
observer.progress(done, done);
Ok(done)
}
/// One tiled pass: every unknown pixel regenerated from the tiles that
/// touch it, the rest kept. Returns the tiles run.
#[allow(clippy::too_many_arguments)]
fn fill_once(
rgb: &mut [f32],
width: usize,
height: usize,
known: &[bool],
model: &mut dyn Inpainter,
stride: usize,
mirror_depth: usize,
mut progress: impl FnMut(usize, usize),
) -> Result<usize, PanoError> {
let t = model.tile();
if t == 0 || known.iter().all(|&k| k) {
return Ok(0);
}
// The padded canvas with mirrored context, and the hole within it.
let ctx = MirroredContext::build(rgb, width, height, known, mirror_depth);
let (pw, ph) = (ctx.width, ctx.height);
// Tiles that touch the hole, on a grid that reaches both far edges.
let starts = |n: usize| -> Vec<usize> {
if n <= t {
return vec![0];
}
let mut v: Vec<usize> = (0..=n - t).step_by(stride.clamp(1, t)).collect();
if *v.last().unwrap_or(&0) != n - t {
v.push(n - t);
}
v
};
let ys = starts(ph);
let xs = starts(pw);
let mut tiles = Vec::new();
for &y in &ys {
for &x in &xs {
if y + t > ph || x + t > pw {
continue;
}
let touches =
(y..y + t).any(|yy| ctx.hole[yy * pw + x..yy * pw + x + t].iter().any(|&h| h));
if touches {
tiles.push((x, y));
}
}
}
// Raised-cosine window, so overlapping tiles blend.
let hann: Vec<f32> = (0..t)
.map(|i| {
let s = ((i as f32 + 1.0) / (t as f32 + 1.0) * std::f32::consts::PI).sin();
s * s + 1e-3
})
.collect();
let mut acc = vec![0.0f32; pw * ph * 3];
let mut wsum = vec![0.0f32; pw * ph];
let mut tile_rgb = vec![0.0f32; t * t * 3];
let mut tile_known = vec![false; t * t];
let total = tiles.len();
for (n, &(x, y)) in tiles.iter().enumerate() {
progress(n, total);
for r in 0..t {
let src = ((y + r) * pw + x) * 3;
tile_rgb[r * t * 3..(r + 1) * t * 3].copy_from_slice(&ctx.rgb[src..src + t * 3]);
let ks = (y + r) * pw + x;
for c in 0..t {
tile_known[r * t + c] = !ctx.hole[ks + c];
}
}
let out = model.fill(&tile_rgb, &tile_known)?;
if out.len() != t * t * 3 {
return Err(PanoError::Model(format!(
"the inpainter returned {} values for a {t}×{t} tile",
out.len()
)));
}
for r in 0..t {
for c in 0..t {
let w = hann[r] * hann[c];
let p = (y + r) * pw + (x + c);
for ch in 0..3 {
acc[p * 3 + ch] += out[(r * t + c) * 3 + ch] * w;
}
wsum[p] += w;
}
}
}
progress(total, total);
// Back into the picture: only the unknown pixels change.
for yy in 0..height {
for xx in 0..width {
let i = yy * width + xx;
if known[i] {
continue;
}
let p = (yy + RING) * pw + (xx + RING);
if wsum[p] > 0.0 {
for ch in 0..3 {
rgb[i * 3 + ch] = (acc[p * 3 + ch] / wsum[p]).clamp(0.0, 1.0);
}
}
}
}
Ok(total)
}
/// Shrink `known` by `iterations` pixels on every side, in place.
///
/// The merge's coverage edge carries a fringe — the last partly-covered
/// pixels of a frame, and whatever the renderer did at the boundary — and
/// a fill that stops exactly at the coverage bit leaves it as a dark line
/// along the seam. Eight pixels at a quarter of the composite's resolution
/// was what it took on the fixture.
pub fn erode(known: &mut [bool], width: usize, height: usize, iterations: usize) {
let mut next = known.to_vec();
for _ in 0..iterations {
for y in 0..height {
for x in 0..width {
let i = y * width + x;
if !known[i] {
continue;
}
let edge = x == 0
|| y == 0
|| x + 1 == width
|| y + 1 == height
|| !known[i - 1]
|| !known[i + 1]
|| !known[i - width]
|| !known[i + width];
next[i] = !edge;
}
}
known.copy_from_slice(&next);
}
}
/// Distance from each pixel to the nearest known one, by two chamfer
/// sweeps — within a few percent of Euclidean, and enough to cut bands.
fn distance_to_known(known: &[bool], width: usize, height: usize) -> Vec<f32> {
let inf = (width + height) as f32;
let mut d: Vec<f32> = known.iter().map(|&k| if k { 0.0 } else { inf }).collect();
let (a, b) = (1.0f32, std::f32::consts::SQRT_2);
for y in 0..height {
for x in 0..width {
let i = y * width + x;
let mut v = d[i];
if x > 0 {
v = v.min(d[i - 1] + a);
}
if y > 0 {
v = v.min(d[i - width] + a);
if x > 0 {
v = v.min(d[i - width - 1] + b);
}
if x + 1 < width {
v = v.min(d[i - width + 1] + b);
}
}
d[i] = v;
}
}
for y in (0..height).rev() {
for x in (0..width).rev() {
let i = y * width + x;
let mut v = d[i];
if x + 1 < width {
v = v.min(d[i + 1] + a);
}
if y + 1 < height {
v = v.min(d[i + width] + a);
if x + 1 < width {
v = v.min(d[i + width + 1] + b);
}
if x > 0 {
v = v.min(d[i + width - 1] + b);
}
}
d[i] = v;
}
}
d
}
/// Distance beyond the edge to distance inside it, folded within `depth`
/// ([`Params::mirror_depth`]): a triangle wave, so the band is read
/// forward and back rather than clamped to one row.
fn fold(d: usize, depth: usize) -> usize {
let period = 2 * depth;
let r = d % period;
if r <= depth {
r
} else {
period - r
}
}
/// The picture on a canvas `RING` wider on every side, with the hole and
/// the ring filled by mirroring the known content across the coverage
/// edge — the nearest `depth` of it, folded — and the hole, the
/// original unknown and nothing else, marked.
struct MirroredContext {
width: usize,
height: usize,
rgb: Vec<f32>,
hole: Vec<bool>,
}
impl MirroredContext {
fn build(rgb: &[f32], width: usize, height: usize, known: &[bool], depth: usize) -> Self {
let fold = |d: usize| fold(d, depth);
let (pw, ph) = (width + 2 * RING, height + 2 * RING);
let mut canvas = vec![0.0f32; pw * ph * 3];
let mut kn = vec![false; pw * ph];
let mut hole = vec![false; pw * ph];
for y in 0..height {
for x in 0..width {
let i = y * width + x;
let p = (y + RING) * pw + (x + RING);
canvas[p * 3..p * 3 + 3].copy_from_slice(&rgb[i * 3..i * 3 + 3]);
kn[p] = known[i];
hole[p] = !known[i];
}
}
// Per column: mirror across the first and last known row.
for x in 0..pw {
let first = (0..ph).find(|&y| kn[y * pw + x]);
let Some(first) = first else { continue };
let last = (0..ph).rev().find(|&y| kn[y * pw + x]).unwrap_or(first);
for y in 0..first {
let m = (first + fold(first - y)).min(last);
let (d, s) = ((y * pw + x) * 3, (m * pw + x) * 3);
canvas.copy_within(s..s + 3, d);
}
for y in last + 1..ph {
let m = last.saturating_sub(fold(y - last)).max(first);
let (d, s) = ((y * pw + x) * 3, (m * pw + x) * 3);
canvas.copy_within(s..s + 3, d);
}
}
// Per row, for the sides, over what is there now.
for y in 0..ph {
let first = (0..pw).find(|&x| kn[y * pw + x]);
let Some(first) = first else { continue };
let last = (0..pw).rev().find(|&x| kn[y * pw + x]).unwrap_or(first);
for x in 0..first {
let m = (first + fold(first - x)).min(last);
let (d, s) = ((y * pw + x) * 3, (y * pw + m) * 3);
canvas.copy_within(s..s + 3, d);
}
for x in last + 1..pw {
let m = last.saturating_sub(fold(x - last)).max(first);
let (d, s) = ((y * pw + x) * 3, (y * pw + m) * 3);
canvas.copy_within(s..s + 3, d);
}
}
MirroredContext {
width: pw,
height: ph,
rgb: canvas,
hole,
}
}
}
#[cfg(test)]
mod tests {
use super::*;
/// The tests' small pictures: a 48-px stride, a given feather.
fn test_params(feather: usize) -> Params {
Params {
stride: 48,
feather,
..Params::default()
}
}
/// Paints every unknown pixel a fixed grey and copies the known ones,
/// and remembers what it was shown.
struct Flat {
tile: usize,
seen: Vec<(Vec<f32>, Vec<bool>)>,
}
impl Inpainter for Flat {
fn tile(&self) -> usize {
self.tile
}
fn fill(&mut self, rgb: &[f32], known: &[bool]) -> Result<Vec<f32>, PanoError> {
self.seen.push((rgb.to_vec(), known.to_vec()));
Ok(rgb
.chunks_exact(3)
.zip(known)
.flat_map(|(p, &k)| if k { [p[0], p[1], p[2]] } else { [0.5; 3] })
.collect())
}
}
fn picture(w: usize, h: usize, border: usize) -> (Vec<f32>, Vec<bool>) {
let mut rgb = vec![0.0; w * h * 3];
let mut known = vec![false; w * h];
for y in 0..h {
for x in 0..w {
let i = y * w + x;
if y >= border && y < h - border {
known[i] = true;
rgb[i * 3] = x as f32 / w as f32;
rgb[i * 3 + 1] = y as f32 / h as f32;
rgb[i * 3 + 2] = 0.25;
}
}
}
(rgb, known)
}
#[test]
fn unknown_pixels_take_the_model_and_known_ones_do_not_move() {
let (mut rgb, known) = picture(300, 200, 20);
let before = rgb.clone();
let mut model = Flat {
tile: 64,
seen: Vec::new(),
};
let tiles = fill_border(
&mut rgb,
300,
200,
&known,
&mut model,
test_params(0),
&mut |_, _| {},
)
.unwrap();
assert!(tiles > 0);
for i in 0..300 * 200 {
if known[i] {
assert_eq!(&rgb[i * 3..i * 3 + 3], &before[i * 3..i * 3 + 3]);
} else {
for c in 0..3 {
assert!((rgb[i * 3 + c] - 0.5).abs() < 1e-4, "pixel {i}");
}
}
}
}
#[test]
fn the_model_is_shown_mirrored_context_not_black() {
let (mut rgb, known) = picture(300, 200, 20);
let mut model = Flat {
tile: 64,
seen: Vec::new(),
};
fill_border(
&mut rgb,
300,
200,
&known,
&mut model,
test_params(0),
&mut |_, _| {},
)
.unwrap();
for (tile_rgb, tile_known) in &model.seen {
let known_non_black = tile_rgb
.chunks_exact(3)
.zip(tile_known)
.filter(|(_, &k)| k)
.any(|(p, _)| p.iter().any(|v| *v > 0.0));
assert!(known_non_black);
}
}
#[test]
fn the_fine_passes_run_in_bands_after_the_coarse_one() {
// A 150-tall hole above and below a picture: the coarse pass sees
// it at a quarter; the fine passes need two bands of BAND pixels.
let (mut rgb, known) = picture(200, 500, 150);
let mut model = Flat {
tile: 64,
seen: Vec::new(),
};
fill_border(
&mut rgb,
200,
500,
&known,
&mut model,
test_params(0),
&mut |_, _| {},
)
.unwrap();
assert!(model.seen.len() > 4);
// Every unknown pixel was reached.
for i in 0..200 * 500 {
if !known[i] {
assert!((rgb[i * 3] - 0.5).abs() < 1e-4, "pixel {i}");
}
}
}
#[test]
fn the_seam_ramps_from_real_to_invented_across_the_feather() {
let (mut rgb, known) = picture(300, 200, 20);
let before = rgb.clone();
let mut model = Flat {
tile: 64,
seen: Vec::new(),
};
fill_border(
&mut rgb,
300,
200,
&known,
&mut model,
test_params(8),
&mut |_, _| {},
)
.unwrap();
// Row 20 is the real edge; the margin runs to row 27. At the edge
// the value is the model's grey, eight rows in it is the picture's.
let at = |y: usize| rgb[(y * 300 + 150) * 3 + 2];
assert!((at(20) - 0.5).abs() < 0.05, "{}", at(20));
assert!((at(29) - before[(29 * 300 + 150) * 3 + 2]).abs() < 1e-4);
let (lo, hi) = (at(20).min(at(29)), at(20).max(at(29)));
assert!(
at(23) > lo + 0.02 && at(23) < hi - 0.02,
"{} between {lo} and {hi}",
at(23)
);
// The hole itself is the model's.
assert!((at(5) - 0.5).abs() < 1e-4);
}
#[test]
fn erosion_shrinks_the_known_region_from_every_edge() {
let (_, mut known) = picture(20, 20, 4);
erode(&mut known, 20, 20, 2);
assert!(known[8 * 20 + 10]);
assert!(!known[5 * 20 + 10]);
assert!(!known[8 * 20 + 1]);
}
#[test]
fn distance_counts_pixels_from_the_known_region() {
let (_, known) = picture(20, 20, 4);
let d = distance_to_known(&known, 20, 20);
assert_eq!(d[4 * 20 + 10], 0.0);
assert!((d[3 * 20 + 10] - 1.0).abs() < 1e-6);
assert!((d[10] - 4.0).abs() < 1e-6);
}
#[test]
fn a_fully_covered_picture_runs_nothing() {
let (mut rgb, known) = picture(100, 100, 0);
let mut model = Flat {
tile: 64,
seen: Vec::new(),
};
assert_eq!(
fill_border(
&mut rgb,
100,
100,
&known,
&mut model,
test_params(0),
&mut |_, _| {}
)
.unwrap(),
0
);
}
#[test]
fn the_context_mirrors_the_top_rows_upward() {
let (rgb, known) = picture(40, 30, 5);
let ctx = MirroredContext::build(&rgb, 40, 30, &known, 48);
let x = RING + 10;
let first = RING + 5;
for k in 1..=4 {
let above = ((first - k) * ctx.width + x) * 3;
let mirror = ((first + k) * ctx.width + x) * 3;
assert_eq!(&ctx.rgb[above..above + 3], &ctx.rgb[mirror..mirror + 3]);
}
assert!(ctx.hole[(RING + 2) * ctx.width + x]);
assert!(!ctx.hole[(RING - 2) * ctx.width + x]);
}
#[test]
fn the_mirror_reaches_no_deeper_than_its_band() {
// A ridge 200 rows in must not appear in the ring: beyond the band
// the reflection folds back towards the edge rather than on into
// the picture.
let (mut rgb, known) = picture(40, 400, 5);
let ridge = 5 + 200;
for x in 0..40 {
rgb[(ridge * 40 + x) * 3..(ridge * 40 + x) * 3 + 3].copy_from_slice(&[0.9, 0.1, 0.1]);
}
let ctx = MirroredContext::build(&rgb, 40, 400, &known, 48);
let x = RING + 10;
for y in 0..RING + 5 {
let p = (y * ctx.width + x) * 3;
assert!(
ctx.rgb[p] < 0.5,
"row {y} of the ring shows the ridge ({:?})",
&ctx.rgb[p..p + 3]
);
}
assert_eq!(fold(0, 48), 0);
assert_eq!(fold(48, 48), 48);
assert_eq!(fold(58, 48), 38);
assert_eq!(fold(96, 48), 0);
assert_eq!(fold(99, 48), 3);
}
}
+367
View File
@@ -0,0 +1,367 @@
//! Pairwise geometry: a homography between two frames, found robustly.
//!
//! Two frames of a panorama are related by a rotation, and a rotation seen
//! through one lens is a homography of the image plane — `H = K R Kᵀ⁻¹`. The
//! homography is estimated first, from matches, because it does not need
//! the focal length; the focal length is then *read off* it (§ below), and
//! the rotation follows from both. This is the order Brown & Lowe (2007)
//! and OpenCV's stitcher use, and it is what makes the pipeline work when
//! EXIF says nothing about the lens.
//!
//! Coordinates throughout are **centred**: the principal point is the
//! origin. The focal formulae assume it, and centring before the DLT also
//! conditions the linear system — Hartley's normalisation, done once by the
//! caller rather than inside every solve.
use crate::linalg::{DMat, Mat3, Vec3};
/// A point in one image, centred on the principal point.
pub type Point = (f64, f64);
/// Apply a homography to a point.
pub fn apply(h: &Mat3, p: Point) -> Option<Point> {
let v = *h * Vec3::new(p.0, p.1, 1.0);
if v.z().abs() < 1e-12 {
return None;
}
Some((v.x() / v.z(), v.y() / v.z()))
}
/// Least-squares homography from at least four correspondences by the
/// direct linear transform, with `h33` fixed at 1.
///
/// Fixing `h33` turns the homogeneous 8×9 system into an ordinary 8-unknown
/// least-squares problem that the normal equations and a Cholesky
/// factorisation solve without an SVD. The one homography it cannot
/// represent — `h33 = 0`, a point at the origin mapped to infinity — does
/// not occur between overlapping frames of one scene.
///
/// The points should be scaled to order one (divide by the focal length or
/// the image size) before calling: the normal equations square the
/// conditioning, and pixel coordinates in the thousands make them singular
/// in `f64`.
pub fn dlt(pairs: &[(Point, Point)]) -> Option<Mat3> {
if pairs.len() < 4 {
return None;
}
// Each pair gives two rows of A h = b with h = (h11..h32).
// x' = (h11 x + h12 y + h13) / (h31 x + h32 y + 1)
// → h11 x + h12 y + h13 - h31 x x' - h32 y x' = x'
let mut ata = DMat::zeros(8);
let mut atb = [0.0f64; 8];
for &((x, y), (xp, yp)) in pairs {
let rows: [([f64; 8], f64); 2] = [
([x, y, 1.0, 0.0, 0.0, 0.0, -x * xp, -y * xp], xp),
([0.0, 0.0, 0.0, x, y, 1.0, -x * yp, -y * yp], yp),
];
for (a, b) in rows {
for i in 0..8 {
atb[i] += a[i] * b;
for j in 0..8 {
ata[(i, j)] += a[i] * a[j];
}
}
}
}
let h = ata.solve_spd(&atb)?;
Some(Mat3([
[h[0], h[1], h[2]],
[h[3], h[4], h[5]],
[h[6], h[7], 1.0],
]))
}
/// A homography with the correspondences that agree with it.
#[derive(Debug, Clone, PartialEq)]
pub struct RobustHomography {
pub h: Mat3,
/// Indices into the input pairs.
pub inliers: Vec<usize>,
}
/// RANSAC over [`dlt`] on four-point samples, then a final least-squares
/// fit over every inlier.
///
/// `threshold` is the reprojection distance, in the same units as the
/// points, within which a pair counts as agreeing. The iteration count
/// adapts to the inlier ratio found so far in the usual way, capped at
/// `max_iterations`. `seed` makes a run reproducible (NFR-MRG-2): the
/// sampling is a small linear congruential generator, not the system's.
pub fn ransac_homography(
pairs: &[(Point, Point)],
threshold: f64,
max_iterations: usize,
seed: u64,
) -> Option<RobustHomography> {
if pairs.len() < 4 {
return None;
}
let n = pairs.len();
let thr2 = threshold * threshold;
let mut rng = Lcg(seed);
let mut best: Option<(Vec<usize>, Mat3)> = None;
let mut iterations = max_iterations;
let mut i = 0;
while i < iterations {
i += 1;
let sample = rng.distinct4(n);
let Some(h) = dlt(&sample.map(|k| pairs[k])) else {
continue;
};
let inliers: Vec<usize> = (0..n).filter(|&k| agrees(&h, pairs[k], thr2)).collect();
if best.as_ref().is_none_or(|(b, _)| inliers.len() > b.len()) {
// Adapt: enough iterations to have drawn one all-inlier sample
// with probability 0.99, given the ratio seen so far.
let w = inliers.len() as f64 / n as f64;
let p_all = w.powi(4);
if p_all > 0.0 && p_all < 1.0 {
let needed = ((1.0 - 0.99f64).ln() / (1.0 - p_all).ln()).ceil() as usize;
iterations = iterations.min(needed.max(i + 1));
}
best = Some((inliers, h));
}
}
let (inliers, h) = best?;
if inliers.len() < 4 {
return None;
}
// Refit on every inlier, and keep the refit only if it did not lose
// support — a least-squares fit over a set with a few borderline points
// can be pulled off the consensus the sample found.
let refit: Vec<(Point, Point)> = inliers.iter().map(|&k| pairs[k]).collect();
let h = match dlt(&refit) {
Some(r) => {
let count = (0..n).filter(|&k| agrees(&r, pairs[k], thr2)).count();
if count >= inliers.len() {
r
} else {
h
}
}
None => h,
};
let inliers: Vec<usize> = (0..n).filter(|&k| agrees(&h, pairs[k], thr2)).collect();
Some(RobustHomography { h, inliers })
}
fn agrees(h: &Mat3, (p, q): (Point, Point), thr2: f64) -> bool {
match apply(h, p) {
Some((x, y)) => {
let (dx, dy) = (x - q.0, y - q.1);
dx * dx + dy * dy <= thr2
}
None => false,
}
}
/// The focal length a homography implies, if it implies one.
///
/// For `H = K R K⁻¹` with `K = diag(f, f, 1)` and the principal point at the
/// origin, the orthonormality of `R` gives two independent estimates of `f²`
/// from the first two rows and two from the first two columns; each is
/// taken where it is positive and the better-conditioned of the pair is
/// chosen, as OpenCV's `focalsFromHomography` does. The geometric mean of
/// the row and column estimates is returned. `None` when the homography is
/// too close to a pure translation to say anything — every estimate is then
/// a ratio of small numbers.
pub fn focal_from_homography(h: &Mat3) -> Option<f64> {
let m = h.0;
let (h00, h01, h02) = (m[0][0], m[0][1], m[0][2]);
let (h10, h11, h12) = (m[1][0], m[1][1], m[1][2]);
let (h20, h21) = (m[2][0], m[2][1]);
let pick = |mut v1: f64, mut v2: f64, d1: f64, d2: f64| -> Option<f64> {
if v1 < v2 {
std::mem::swap(&mut v1, &mut v2);
}
if v1 > 0.0 && v2 > 0.0 {
Some((if d1.abs() > d2.abs() { v1 } else { v2 }).sqrt())
} else if v1 > 0.0 {
Some(v1.sqrt())
} else {
None
}
};
// From the third row.
let d1 = h20 * h21;
let d2 = (h21 - h20) * (h21 + h20);
let f1 = if d1.abs() > 1e-12 || d2.abs() > 1e-12 {
let v1 = if d1.abs() > 1e-12 {
-(h00 * h01 + h10 * h11) / d1
} else {
f64::NAN
};
let v2 = if d2.abs() > 1e-12 {
(h00 * h00 + h10 * h10 - h01 * h01 - h11 * h11) / d2
} else {
f64::NAN
};
pick(nan_to_neg(v1), nan_to_neg(v2), d1, d2)
} else {
None
};
// From the third column.
let d1 = h00 * h10 + h01 * h11;
let d2 = h00 * h00 + h01 * h01 - h10 * h10 - h11 * h11;
let f0 = if d1.abs() > 1e-12 || d2.abs() > 1e-12 {
let v1 = if d1.abs() > 1e-12 {
-h02 * h12 / d1
} else {
f64::NAN
};
let v2 = if d2.abs() > 1e-12 {
(h12 * h12 - h02 * h02) / d2
} else {
f64::NAN
};
pick(nan_to_neg(v1), nan_to_neg(v2), d1, d2)
} else {
None
};
match (f0, f1) {
(Some(a), Some(b)) => Some((a * b).sqrt()),
(Some(a), None) | (None, Some(a)) => Some(a),
(None, None) => None,
}
}
fn nan_to_neg(v: f64) -> f64 {
if v.is_finite() {
v
} else {
-1.0
}
}
/// The rotation a homography encodes for a known focal length:
/// `R = K⁻¹ H K`, re-orthonormalised, with the scale of `H` divided out.
pub fn rotation_from_homography(h: &Mat3, f: f64) -> Mat3 {
let m = h.0;
// K⁻¹ H K with K = diag(f, f, 1): scale the third row by f and the
// third column by 1/f.
let r = Mat3([
[m[0][0], m[0][1], m[0][2] / f],
[m[1][0], m[1][1], m[1][2] / f],
[m[2][0] * f, m[2][1] * f, m[2][2]],
]);
r.orthonormalised()
}
/// A small deterministic generator for RANSAC's samples.
struct Lcg(u64);
impl Lcg {
fn next(&mut self) -> u64 {
// Knuth's MMIX constants.
self.0 = self
.0
.wrapping_mul(6364136223846793005)
.wrapping_add(1442695040888963407);
self.0 >> 33
}
fn below(&mut self, n: usize) -> usize {
(self.next() % n as u64) as usize
}
fn distinct4(&mut self, n: usize) -> [usize; 4] {
let mut s = [0usize; 4];
for i in 0..4 {
loop {
let k = self.below(n);
if !s[..i].contains(&k) {
s[i] = k;
break;
}
}
}
s
}
}
#[cfg(test)]
mod tests {
use super::*;
/// Points under a known rotation seen through a known focal length,
/// in centred image coordinates scaled by that focal length.
fn synthetic(f: f64, r: Mat3, n: usize, noise: f64, seed: u64) -> Vec<(Point, Point)> {
let mut rng = Lcg(seed);
let mut out = Vec::new();
while out.len() < n {
// A point on the first image plane, within ±0.3 f of centre.
let x = (rng.below(6001) as f64 - 3000.0) / 10000.0;
let y = (rng.below(4001) as f64 - 2000.0) / 10000.0;
let b = Vec3::new(x, y, 1.0);
let v = r * b;
if v.z() <= 0.2 {
continue;
}
let nx = (rng.below(2001) as f64 - 1000.0) / 1000.0 * noise;
let ny = (rng.below(2001) as f64 - 1000.0) / 1000.0 * noise;
out.push(((x, y), (v.x() / v.z() + nx, v.y() / v.z() + ny)));
}
let _ = f;
out
}
#[test]
fn dlt_recovers_a_known_homography_exactly() {
let r = Mat3::exp(Vec3::new(0.05, 0.3, 0.02));
let pairs = synthetic(1.0, r, 12, 0.0, 1);
let h = dlt(&pairs).expect("solvable");
for &(p, q) in &pairs {
let (x, y) = apply(&h, p).unwrap();
assert!((x - q.0).abs() < 1e-9 && (y - q.1).abs() < 1e-9);
}
}
#[test]
fn ransac_finds_the_consensus_among_outliers() {
let r = Mat3::exp(Vec3::new(-0.02, 0.25, 0.01));
let mut pairs = synthetic(1.0, r, 60, 0.0005, 2);
// Forty outliers: wrong second point.
let mut rng = Lcg(9);
for _ in 0..40 {
let k = rng.below(60);
let (p, _) = pairs[k];
pairs.push((p, ((rng.below(1000) as f64 - 500.0) / 1000.0, 0.1)));
}
let robust = ransac_homography(&pairs, 0.003, 500, 3).expect("found");
assert!(
robust.inliers.len() >= 55,
"{} inliers",
robust.inliers.len()
);
assert!(robust.inliers.iter().all(|&k| k < 60));
}
#[test]
fn focal_is_read_off_a_rotation_homography() {
// H in *pixel* coordinates for f = 1400: K R K⁻¹.
let f = 1400.0;
let r = Mat3::exp(Vec3::new(0.03, 0.35, -0.01));
let m = r.0;
let h = Mat3([
[m[0][0], m[0][1], m[0][2] * f],
[m[1][0], m[1][1], m[1][2] * f],
[m[2][0] / f, m[2][1] / f, m[2][2]],
]);
let est = focal_from_homography(&h).expect("estimable");
assert!((est - f).abs() / f < 1e-6, "{est}");
let back = rotation_from_homography(&h, f);
for (row, truth) in back.0.iter().zip(&m) {
for (a, b) in row.iter().zip(truth) {
assert!((a - b).abs() < 1e-9);
}
}
}
#[test]
fn the_identity_implies_no_focal() {
assert!(focal_from_homography(&Mat3::IDENTITY).is_none());
}
}
+262
View File
@@ -0,0 +1,262 @@
//! The grayscale proxy a detector reads.
//!
//! Alignment runs on proxies (FR-MRG-7) — a detector at 1024 px sees
//! everything it needs, and the full-resolution frames never leave the GPU.
//! This is that proxy: one channel, `f32` in `0.0..=1.0`, upright, and no
//! larger than the detector's fixed input.
/// A single-channel image, row-major, values in `0.0..=1.0`.
#[derive(Debug, Clone, PartialEq)]
pub struct Gray {
pub width: usize,
pub height: usize,
pub data: Vec<f32>,
}
impl Gray {
/// From tightly packed 8-bit RGBA, by the Rec. 709 luma weights.
///
/// The proxy is what a detector looks at, not what the photographer
/// sees, so which luma is used matters less than that it is the same one
/// for every frame — a keypoint's descriptor must not change between two
/// frames because they were converted differently.
pub fn from_rgba8(rgba: &[u8], width: usize, height: usize) -> Gray {
let n = width * height;
assert!(
rgba.len() >= n * 4,
"rgba buffer is short for {width}×{height}"
);
let data = rgba[..n * 4]
.chunks_exact(4)
.map(|p| {
(0.2126 * f32::from(p[0]) + 0.7152 * f32::from(p[1]) + 0.0722 * f32::from(p[2]))
/ 255.0
})
.collect();
Gray {
width,
height,
data,
}
}
/// Apply an EXIF orientation so the image is upright.
///
/// Learned detectors are not rotation-invariant — a descriptor of a
/// feature seen sideways is a different descriptor — and a portrait set
/// (the 6D fixture is one) would match poorly or not at all fed as
/// stored. The camera says which way is up; the proxy is turned before
/// anything looks at it, and the composite is written upright.
///
/// The value is the EXIF `Orientation` tag. Mirrored values (2, 4, 5, 7)
/// are not produced by any camera and are treated as their unmirrored
/// counterparts.
pub fn oriented(&self, orientation: u16) -> Gray {
match orientation {
3 | 4 => self.rotated_180(),
6 | 5 => self.rotated_90_cw(),
8 | 7 => self.rotated_90_ccw(),
_ => self.clone(),
}
}
fn rotated_90_cw(&self) -> Gray {
let (w, h) = (self.width, self.height);
let mut data = vec![0.0; w * h];
for y in 0..h {
for x in 0..w {
// Source (x, y) lands at (h - 1 - y, x) in an h-wide image.
data[x * h + (h - 1 - y)] = self.data[y * w + x];
}
}
Gray {
width: h,
height: w,
data,
}
}
fn rotated_90_ccw(&self) -> Gray {
let (w, h) = (self.width, self.height);
let mut data = vec![0.0; w * h];
for y in 0..h {
for x in 0..w {
// Source (x, y) lands at (y, w - 1 - x) in an h-wide image.
data[(w - 1 - x) * h + y] = self.data[y * w + x];
}
}
Gray {
width: h,
height: w,
data,
}
}
fn rotated_180(&self) -> Gray {
let mut data = self.data.clone();
data.reverse();
Gray {
width: self.width,
height: self.height,
data,
}
}
/// Resample to exactly `width × height` by area averaging on the way
/// down and bilinear on the way up.
///
/// Area averaging, not point sampling, for a reduction: a 5472 px frame
/// to 1024 is a factor of five, and picking one source pixel in
/// twenty-five aliases every edge the detector is looking for.
pub fn resampled(&self, width: usize, height: usize) -> Gray {
if width == self.width && height == self.height {
return self.clone();
}
let mut data = vec![0.0f32; width * height];
let sx = self.width as f64 / width as f64;
let sy = self.height as f64 / height as f64;
if sx >= 1.0 && sy >= 1.0 {
for oy in 0..height {
let y0 = (oy as f64 * sy) as usize;
let y1 = (((oy + 1) as f64 * sy) as usize).clamp(y0 + 1, self.height);
for ox in 0..width {
let x0 = (ox as f64 * sx) as usize;
let x1 = (((ox + 1) as f64 * sx) as usize).clamp(x0 + 1, self.width);
let mut sum = 0.0f32;
for y in y0..y1 {
let row = &self.data[y * self.width..(y + 1) * self.width];
sum += row[x0..x1].iter().sum::<f32>();
}
data[oy * width + ox] = sum / ((y1 - y0) * (x1 - x0)) as f32;
}
}
} else {
for oy in 0..height {
let fy = ((oy as f64 + 0.5) * sy - 0.5).max(0.0);
let y0 = (fy as usize).min(self.height - 1);
let y1 = (y0 + 1).min(self.height - 1);
let ty = (fy - y0 as f64) as f32;
for ox in 0..width {
let fx = ((ox as f64 + 0.5) * sx - 0.5).max(0.0);
let x0 = (fx as usize).min(self.width - 1);
let x1 = (x0 + 1).min(self.width - 1);
let tx = (fx - x0 as f64) as f32;
let p = |x: usize, y: usize| self.data[y * self.width + x];
let top = p(x0, y0) * (1.0 - tx) + p(x1, y0) * tx;
let bot = p(x0, y1) * (1.0 - tx) + p(x1, y1) * tx;
data[oy * width + ox] = top * (1.0 - ty) + bot * ty;
}
}
}
Gray {
width,
height,
data,
}
}
/// Scale so the image fits inside `max_width × max_height`, preserving
/// aspect, never enlarging. Returns the image and the scale applied,
/// which is what maps a proxy keypoint back to the source.
pub fn fitted(&self, max_width: usize, max_height: usize) -> (Gray, f64) {
let scale = (max_width as f64 / self.width as f64)
.min(max_height as f64 / self.height as f64)
.min(1.0);
let w = ((self.width as f64 * scale).round() as usize).max(1);
let h = ((self.height as f64 * scale).round() as usize).max(1);
(self.resampled(w, h), w as f64 / self.width as f64)
}
/// Copy into the top-left of a `width × height` canvas, zero elsewhere.
///
/// The detector's input is a fixed shape (S15.2), and a frame that fits
/// inside it is padded rather than stretched: stretching changes the
/// aspect and with it every descriptor.
pub fn padded(&self, width: usize, height: usize) -> Gray {
assert!(self.width <= width && self.height <= height);
let mut data = vec![0.0; width * height];
for y in 0..self.height {
data[y * width..y * width + self.width]
.copy_from_slice(&self.data[y * self.width..(y + 1) * self.width]);
}
Gray {
width,
height,
data,
}
}
}
#[cfg(test)]
mod tests {
use super::*;
fn ramp(w: usize, h: usize) -> Gray {
Gray {
width: w,
height: h,
data: (0..w * h).map(|i| i as f32).collect(),
}
}
#[test]
fn rotating_four_quarter_turns_is_the_identity() {
let g = ramp(5, 3);
let mut r = g.clone();
for _ in 0..4 {
r = r.rotated_90_cw();
}
assert_eq!(r, g);
assert_eq!(g.rotated_90_cw().rotated_90_ccw(), g);
assert_eq!(g.rotated_180().rotated_180(), g);
}
#[test]
fn a_clockwise_turn_moves_the_top_left_to_the_top_right() {
// 2×3 image, pixel values by position.
let g = ramp(2, 3);
let r = g.rotated_90_cw();
assert_eq!((r.width, r.height), (3, 2));
// Top-left of source (value 0) is at top-right of result.
assert_eq!(r.data[2], 0.0);
// Bottom-left of source (value 4) is at top-left of result.
assert_eq!(r.data[0], 4.0);
}
#[test]
fn orientation_8_is_a_counter_clockwise_turn() {
let g = ramp(4, 2);
assert_eq!(g.oriented(8), g.rotated_90_ccw());
assert_eq!(g.oriented(6), g.rotated_90_cw());
assert_eq!(g.oriented(1), g);
}
#[test]
fn downsampling_by_two_averages_blocks() {
let g = Gray {
width: 4,
height: 2,
data: vec![0.0, 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0],
};
let r = g.resampled(2, 1);
assert_eq!(r.data, vec![2.5, 4.5]);
}
#[test]
fn fitting_never_enlarges_and_reports_the_scale() {
let g = ramp(100, 50);
let (f, s) = g.fitted(1024, 768);
assert_eq!((f.width, f.height), (100, 50));
assert_eq!(s, 1.0);
let (f, s) = g.fitted(50, 50);
assert_eq!((f.width, f.height), (50, 25));
assert_eq!(s, 0.5);
}
#[test]
fn padding_places_the_image_at_the_origin() {
let g = ramp(2, 2);
let p = g.padded(3, 3);
assert_eq!(p.data, vec![0.0, 1.0, 0.0, 2.0, 3.0, 0.0, 0.0, 0.0, 0.0]);
}
}
+86
View File
@@ -0,0 +1,86 @@
//! TRACES: FR-MRG-1 | FR-MRG-10
//! Panorama geometry — from several frames to the rotations that relate
//! them, and the projections that lay them out.
//!
//! This is the CPU half of a merge (FR-MRG-10): keypoints, matching, the
//! rotation solve and the choice of output surface. The per-pixel half —
//! rendering, warping, seams, blending — is the GPU's and lives in
//! `dr-gpu`, driven from above; nothing here touches a full-resolution
//! pixel. The split is the whole design (panorama.md §4): everything in
//! this crate is bounded by the number of frames, not the size of the
//! composite, and runs on proxies.
//!
//! # Layout
//!
//! - [`image`] — the grayscale proxy a detector reads: oriented, resampled.
//! - [`features`] — keypoints with descriptors, and the XFeat decoder.
//! - [`xfeat`] — the network under tract (feature `xfeat`).
//! - [`matching`] — mutual nearest neighbours.
//! - [`homography`] — a robust pairwise homography, the focal length read
//! off it, and the rotation it implies.
//! - [`bundle`] — every rotation and the focal length refined together.
//! - [`align`] — the whole thing, from features to cameras, honest about
//! what it could not place.
//! - [`projection`] — perspective, cylindrical, spherical.
//! - [`linalg`] — the small dense algebra all of it uses.
//!
//! # What it depends on
//!
//! Nothing, without the `xfeat` feature: the geometry is pure Rust with
//! hand-rolled linear algebra (`linalg` says why) so that it tests without
//! a model, a GPU or a device, on synthetic sets whose answer is known
//! exactly. With the feature it adds the same `ort`-over-tract runtime the
//! rest of the application already carries.
pub mod align;
pub mod bundle;
pub mod features;
pub mod fill;
pub mod homography;
pub mod image;
pub mod linalg;
pub mod matching;
#[cfg(feature = "xfeat")]
pub mod migan;
pub mod projection;
#[cfg(feature = "xfeat")]
pub mod xfeat;
pub use align::{align, AlignOptions, Alignment, Link, Unaligned};
pub use bundle::Cameras;
pub use features::{Features, Keypoint};
pub use fill::{fill_border, Inpainter, Observer, Params as FillParams};
pub use image::Gray;
pub use projection::Projection;
#[derive(Debug, thiserror::Error)]
pub enum PanoError {
#[error("bad input: {0}")]
Input(String),
#[error("geometry: {0}")]
Geometry(String),
#[error("model: {0}")]
Model(String),
#[error("could not read the model: {0}")]
ModelRead(#[source] std::io::Error),
#[cfg(feature = "xfeat")]
#[error("inference: {0}")]
Inference(#[source] ort::Error),
}
#[cfg(feature = "xfeat")]
impl From<ort::Error> for PanoError {
fn from(e: ort::Error) -> Self {
PanoError::Inference(e)
}
}
#[cfg(feature = "xfeat")]
impl From<dr_inference_engine::Error> for PanoError {
fn from(e: dr_inference_engine::Error) -> Self {
match e {
dr_inference_engine::Error::Inference(e) => PanoError::Inference(e),
dr_inference_engine::Error::Io(e) => PanoError::ModelRead(e),
}
}
}
+371
View File
@@ -0,0 +1,371 @@
//! The small dense linear algebra the geometry needs, and nothing more.
//!
//! Hand-rolled rather than pulled in, and the decision was made on purpose
//! (2026-09-19): the largest system this crate ever solves is a rotation
//! per frame plus one focal length — forty unknowns for a dozen frames —
//! and everything else is three-vectors. A general linear-algebra crate
//! would be the largest dependency in `dr-pano` by an order of magnitude,
//! for a Cholesky factorisation that is thirty lines.
//!
//! `f64` throughout. The geometry is solved once per merge on a few thousand
//! matches; there is no reason to give up precision for speed here, and the
//! bundle adjustment's normal equations are poorly conditioned enough near
//! convergence that `f32` would stall it.
use std::ops::{Add, Index, IndexMut, Mul, Neg, Sub};
/// A vector in three dimensions.
#[derive(Debug, Clone, Copy, PartialEq, Default)]
pub struct Vec3(pub [f64; 3]);
impl Vec3 {
pub const fn new(x: f64, y: f64, z: f64) -> Self {
Vec3([x, y, z])
}
pub fn dot(self, o: Vec3) -> f64 {
self.0[0] * o.0[0] + self.0[1] * o.0[1] + self.0[2] * o.0[2]
}
pub fn cross(self, o: Vec3) -> Vec3 {
Vec3([
self.0[1] * o.0[2] - self.0[2] * o.0[1],
self.0[2] * o.0[0] - self.0[0] * o.0[2],
self.0[0] * o.0[1] - self.0[1] * o.0[0],
])
}
pub fn norm(self) -> f64 {
self.dot(self).sqrt()
}
/// The unit vector along `self`, or `self` unchanged if it is zero.
pub fn normalised(self) -> Vec3 {
let n = self.norm();
if n > 0.0 {
self * (1.0 / n)
} else {
self
}
}
pub fn x(self) -> f64 {
self.0[0]
}
pub fn y(self) -> f64 {
self.0[1]
}
pub fn z(self) -> f64 {
self.0[2]
}
}
impl Add for Vec3 {
type Output = Vec3;
fn add(self, o: Vec3) -> Vec3 {
Vec3([self.0[0] + o.0[0], self.0[1] + o.0[1], self.0[2] + o.0[2]])
}
}
impl Sub for Vec3 {
type Output = Vec3;
fn sub(self, o: Vec3) -> Vec3 {
Vec3([self.0[0] - o.0[0], self.0[1] - o.0[1], self.0[2] - o.0[2]])
}
}
impl Mul<f64> for Vec3 {
type Output = Vec3;
fn mul(self, s: f64) -> Vec3 {
Vec3([self.0[0] * s, self.0[1] * s, self.0[2] * s])
}
}
impl Neg for Vec3 {
type Output = Vec3;
fn neg(self) -> Vec3 {
Vec3([-self.0[0], -self.0[1], -self.0[2]])
}
}
/// A 3×3 matrix, row-major.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Mat3(pub [[f64; 3]; 3]);
impl Mat3 {
pub const IDENTITY: Mat3 = Mat3([[1.0, 0.0, 0.0], [0.0, 1.0, 0.0], [0.0, 0.0, 1.0]]);
/// The matrix whose columns are `a`, `b`, `c`.
pub fn from_columns(a: Vec3, b: Vec3, c: Vec3) -> Mat3 {
Mat3([
[a.0[0], b.0[0], c.0[0]],
[a.0[1], b.0[1], c.0[1]],
[a.0[2], b.0[2], c.0[2]],
])
}
pub fn transpose(self) -> Mat3 {
let m = self.0;
Mat3([
[m[0][0], m[1][0], m[2][0]],
[m[0][1], m[1][1], m[2][1]],
[m[0][2], m[1][2], m[2][2]],
])
}
pub fn column(self, i: usize) -> Vec3 {
Vec3([self.0[0][i], self.0[1][i], self.0[2][i]])
}
pub fn trace(self) -> f64 {
self.0[0][0] + self.0[1][1] + self.0[2][2]
}
/// The rotation about `axis` (any length) by `angle` radians — Rodrigues.
pub fn rotation(axis: Vec3, angle: f64) -> Mat3 {
let k = axis.normalised();
let (s, c) = angle.sin_cos();
let t = 1.0 - c;
let (x, y, z) = (k.0[0], k.0[1], k.0[2]);
Mat3([
[t * x * x + c, t * x * y - s * z, t * x * z + s * y],
[t * x * y + s * z, t * y * y + c, t * y * z - s * x],
[t * x * z - s * y, t * y * z + s * x, t * z * z + c],
])
}
/// The rotation whose axis-angle vector is `w` (direction is the axis,
/// length is the angle). The exponential map; [`Self::log`] inverts it.
pub fn exp(w: Vec3) -> Mat3 {
let angle = w.norm();
if angle < 1e-12 {
// First-order: I + [w]×, which is what the limit is and avoids
// dividing by the angle.
let (x, y, z) = (w.0[0], w.0[1], w.0[2]);
return Mat3([[1.0, -z, y], [z, 1.0, -x], [-y, x, 1.0]]);
}
Mat3::rotation(w, angle)
}
/// The axis-angle vector of a rotation matrix. Inverse of [`Self::exp`].
pub fn log(self) -> Vec3 {
let m = self.0;
let cos = ((self.trace() - 1.0) * 0.5).clamp(-1.0, 1.0);
let axis = Vec3([m[2][1] - m[1][2], m[0][2] - m[2][0], m[1][0] - m[0][1]]);
if cos > 1.0 - 1e-6 {
// Small angle: `acos` near 1 loses everything below ~1e-8 to
// rounding, but the antisymmetric part is `2 sin θ · axis` and
// keeps it. First order, exact to the precision that matters.
return axis * 0.5;
}
let angle = cos.acos();
if angle > std::f64::consts::PI - 1e-6 {
// Near π the antisymmetric part vanishes; take the axis from the
// symmetric part instead. Rare for a panorama, but the solver may
// pass through it on a bad start and must not return NaN.
let d = Vec3([
((m[0][0] + 1.0) * 0.5).max(0.0).sqrt(),
((m[1][1] + 1.0) * 0.5).max(0.0).sqrt(),
((m[2][2] + 1.0) * 0.5).max(0.0).sqrt(),
]);
return d.normalised() * angle;
}
axis * (angle / (2.0 * angle.sin()))
}
/// Re-orthonormalise a matrix that has drifted from a rotation through
/// accumulated products. Gram–Schmidt on the columns; cheap and adequate
/// for drift of the size floating-point products produce.
pub fn orthonormalised(self) -> Mat3 {
let a = self.column(0).normalised();
let b = (self.column(1) - a * a.dot(self.column(1))).normalised();
let c = a.cross(b);
Mat3::from_columns(a, b, c)
}
}
impl Mul<Vec3> for Mat3 {
type Output = Vec3;
fn mul(self, v: Vec3) -> Vec3 {
let m = self.0;
Vec3([
m[0][0] * v.0[0] + m[0][1] * v.0[1] + m[0][2] * v.0[2],
m[1][0] * v.0[0] + m[1][1] * v.0[1] + m[1][2] * v.0[2],
m[2][0] * v.0[0] + m[2][1] * v.0[1] + m[2][2] * v.0[2],
])
}
}
impl Mul for Mat3 {
type Output = Mat3;
fn mul(self, o: Mat3) -> Mat3 {
let mut r = [[0.0; 3]; 3];
for (i, row) in r.iter_mut().enumerate() {
for (j, cell) in row.iter_mut().enumerate() {
*cell = (0..3).map(|k| self.0[i][k] * o.0[k][j]).sum();
}
}
Mat3(r)
}
}
/// A dense square matrix, for the normal equations.
#[derive(Debug, Clone, PartialEq)]
pub struct DMat {
n: usize,
data: Vec<f64>,
}
impl DMat {
pub fn zeros(n: usize) -> DMat {
DMat {
n,
data: vec![0.0; n * n],
}
}
pub fn n(&self) -> usize {
self.n
}
/// Solve `self · x = b` for a symmetric positive-definite `self` by
/// Cholesky factorisation. `None` if the matrix is not positive definite,
/// which for the normal equations means the problem is not determined by
/// the data — a frame with no matches, for instance — and the caller
/// should say so rather than proceed.
///
/// Destroys neither input: the factor is built in a copy. The systems
/// here are at most a few dozen unknowns and the copy is nothing.
pub fn solve_spd(&self, b: &[f64]) -> Option<Vec<f64>> {
let n = self.n;
debug_assert_eq!(b.len(), n);
let mut l = vec![0.0; n * n];
for j in 0..n {
let mut d = self[(j, j)];
for k in 0..j {
d -= l[j * n + k] * l[j * n + k];
}
if d <= 0.0 || !d.is_finite() {
return None;
}
let djj = d.sqrt();
l[j * n + j] = djj;
for i in j + 1..n {
let mut s = self[(i, j)];
for k in 0..j {
s -= l[i * n + k] * l[j * n + k];
}
l[i * n + j] = s / djj;
}
}
// Forward: L y = b.
let mut y = vec![0.0; n];
for i in 0..n {
let mut s = b[i];
for k in 0..i {
s -= l[i * n + k] * y[k];
}
y[i] = s / l[i * n + i];
}
// Back: Lᵀ x = y.
let mut x = vec![0.0; n];
for i in (0..n).rev() {
let mut s = y[i];
for k in i + 1..n {
s -= l[k * n + i] * x[k];
}
x[i] = s / l[i * n + i];
}
Some(x)
}
}
impl Index<(usize, usize)> for DMat {
type Output = f64;
fn index(&self, (i, j): (usize, usize)) -> &f64 {
&self.data[i * self.n + j]
}
}
impl IndexMut<(usize, usize)> for DMat {
fn index_mut(&mut self, (i, j): (usize, usize)) -> &mut f64 {
&mut self.data[i * self.n + j]
}
}
#[cfg(test)]
mod tests {
use super::*;
fn close(a: f64, b: f64) -> bool {
(a - b).abs() < 1e-9
}
#[test]
fn exp_and_log_are_inverses() {
for w in [
Vec3::new(0.1, -0.2, 0.3),
Vec3::new(1.0, 0.0, 0.0),
Vec3::new(0.0, 0.0, 2.5),
Vec3::new(1e-9, 0.0, 0.0),
] {
let back = Mat3::exp(w).log();
for i in 0..3 {
assert!(close(back.0[i], w.0[i]), "{w:?} -> {back:?}");
}
}
}
#[test]
fn a_rotation_is_orthonormal_and_preserves_length() {
let r = Mat3::exp(Vec3::new(0.4, 0.5, -0.6));
let rt = r.transpose() * r;
for i in 0..3 {
for j in 0..3 {
assert!(close(rt.0[i][j], Mat3::IDENTITY.0[i][j]));
}
}
let v = Vec3::new(1.0, 2.0, 3.0);
assert!(close((r * v).norm(), v.norm()));
}
#[test]
fn rotation_about_z_turns_x_towards_y() {
let r = Mat3::rotation(Vec3::new(0.0, 0.0, 1.0), std::f64::consts::FRAC_PI_2);
let v = r * Vec3::new(1.0, 0.0, 0.0);
assert!(close(v.x(), 0.0) && close(v.y(), 1.0) && close(v.z(), 0.0));
}
#[test]
fn cholesky_solves_a_small_spd_system() {
// A = Bᵀ B for a random-ish B is SPD by construction.
let b = [
[2.0, 1.0, 0.0],
[1.0, 3.0, 1.0],
[0.0, 1.0, 4.0],
[1.0, 1.0, 1.0],
];
let mut a = DMat::zeros(3);
for i in 0..3 {
for j in 0..3 {
a[(i, j)] = (0..4).map(|k| b[k][i] * b[k][j]).sum();
}
}
let x_true = [1.0, -2.0, 0.5];
let rhs: Vec<f64> = (0..3)
.map(|i| (0..3).map(|j| a[(i, j)] * x_true[j]).sum())
.collect();
let x = a.solve_spd(&rhs).expect("spd");
for i in 0..3 {
assert!(close(x[i], x_true[i]), "{x:?}");
}
}
#[test]
fn cholesky_refuses_an_indefinite_matrix() {
let mut a = DMat::zeros(2);
a[(0, 0)] = 1.0;
a[(1, 1)] = -1.0;
assert!(a.solve_spd(&[1.0, 1.0]).is_none());
}
}
+173
View File
@@ -0,0 +1,173 @@
//! Descriptor matching between two images.
//!
//! Mutual nearest neighbour on cosine similarity, with a floor on the
//! similarity — the reference XFeat's own matcher (`match_mkpts`,
//! `min_cossim = 0.82`). For a panorama that is enough: one lens, one
//! scene, near-pure rotation and 20–40 % overlap make the matching problem
//! easy, and what is hard — sky, repeated structure, exposure drift — is
//! handled by the detector's descriptors and by RANSAC downstream, not by a
//! cleverer matcher. A learned matcher (LightGlue) is the step after this
//! one fails on a real set, and it has not (panorama.md §6).
//!
//! Brute force. `4096 × 4096 × 64` multiply-adds is a billion per pair,
//! and a twelve-frame set has sixty-six pairs: a minute single-threaded
//! and scalar (measured 2026-09-19: 51 s), a few seconds vectorised across
//! the cores. Not worth an index, but worth doing properly.
use crate::features::{Features, DESCRIPTOR_LEN};
const _: () = assert!(DESCRIPTOR_LEN.is_multiple_of(8));
/// A correspondence: keypoint `a` in the first image matches keypoint `b`
/// in the second, with the cosine similarity of their descriptors.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Match {
pub a: usize,
pub b: usize,
pub similarity: f32,
}
/// Match two sets of features.
///
/// A pair is kept when each is the other's nearest neighbour and their
/// similarity is at least `min_similarity`.
pub fn match_features(a: &Features, b: &Features, min_similarity: f32) -> Vec<Match> {
if a.is_empty() || b.is_empty() {
return Vec::new();
}
let (na, nb) = (a.len(), b.len());
// The whole similarity matrix, once. Both nearest-neighbour directions
// read it, which halves the multiply-adds against computing each
// direction on its own; 4096 × 4096 × f32 is 64 MB, transient.
let mut sim = vec![0.0f32; na * nb];
let threads = std::thread::available_parallelism()
.map(usize::from)
.unwrap_or(1)
.clamp(1, 16);
let rows_per = na.div_ceil(threads);
std::thread::scope(|scope| {
for (t, chunk) in sim.chunks_mut(rows_per * nb).enumerate() {
scope.spawn(move || {
let first = t * rows_per;
for (r, row) in chunk.chunks_mut(nb).enumerate() {
let da = a.descriptor(first + r);
for (j, cell) in row.iter_mut().enumerate() {
*cell = dot(da, b.descriptor(j));
}
}
});
}
});
// Best in `b` for each `a`, and best in `a` for each `b`.
let best_ab: Vec<(usize, f32)> = sim
.chunks_exact(nb)
.map(|row| {
row.iter().enumerate().fold(
(0usize, f32::MIN),
|acc, (j, &s)| if s > acc.1 { (j, s) } else { acc },
)
})
.collect();
let mut best_ba = vec![(0usize, f32::MIN); nb];
for (i, row) in sim.chunks_exact(nb).enumerate() {
for (j, &s) in row.iter().enumerate() {
if s > best_ba[j].1 {
best_ba[j] = (i, s);
}
}
}
best_ab
.iter()
.enumerate()
.filter_map(|(ia, &(ib, s))| {
(best_ba[ib].0 == ia && s >= min_similarity).then_some(Match {
a: ia,
b: ib,
similarity: s,
})
})
.collect()
}
#[inline]
fn dot(a: &[f32], b: &[f32]) -> f32 {
// Eight independent accumulators over exact 8-lane chunks: the shape
// the compiler turns into one vector multiply-add per chunk, and no
// bounds checks inside the loop. `DESCRIPTOR_LEN` is a multiple of 8.
let (a, b) = (&a[..DESCRIPTOR_LEN], &b[..DESCRIPTOR_LEN]);
let mut acc = [0.0f32; 8];
for (ca, cb) in a.chunks_exact(8).zip(b.chunks_exact(8)) {
for k in 0..8 {
acc[k] += ca[k] * cb[k];
}
}
acc.iter().sum()
}
#[cfg(test)]
mod tests {
use super::*;
use crate::features::Keypoint;
/// Features whose descriptors are unit vectors along the given axes.
fn along(axes: &[usize]) -> Features {
let mut descriptors = vec![0.0; axes.len() * DESCRIPTOR_LEN];
for (i, &ax) in axes.iter().enumerate() {
descriptors[i * DESCRIPTOR_LEN + ax] = 1.0;
}
Features {
keypoints: axes
.iter()
.map(|_| Keypoint {
x: 0.0,
y: 0.0,
score: 1.0,
})
.collect(),
descriptors,
width: 1,
height: 1,
}
}
#[test]
fn identical_descriptors_match_mutually() {
let a = along(&[0, 1, 2]);
let b = along(&[2, 0, 1]);
let m = match_features(&a, &b, 0.8);
let mut pairs: Vec<(usize, usize)> = m.iter().map(|m| (m.a, m.b)).collect();
pairs.sort();
assert_eq!(pairs, vec![(0, 1), (1, 2), (2, 0)]);
assert!(m.iter().all(|m| (m.similarity - 1.0).abs() < 1e-6));
}
#[test]
fn a_descriptor_with_no_counterpart_is_unmatched() {
let a = along(&[0, 1, 5]);
let b = along(&[0, 1]);
let m = match_features(&a, &b, 0.8);
assert_eq!(m.len(), 2);
assert!(m.iter().all(|m| m.a != 2));
}
#[test]
fn mutuality_breaks_a_one_sided_match() {
// b0 is the nearest to both a0 and a1, but a0 is its nearest — a1
// must not be matched to it.
let mut a = along(&[0, 0]);
a.descriptors[DESCRIPTOR_LEN] = 0.9;
a.descriptors[DESCRIPTOR_LEN + 1] = (1.0f32 - 0.81).sqrt();
let b = along(&[0]);
let m = match_features(&a, &b, 0.0);
assert_eq!(m.len(), 1);
assert_eq!((m[0].a, m[0].b), (0, 0));
}
#[test]
fn empty_input_is_empty_output() {
assert!(match_features(&along(&[]), &along(&[1]), 0.5).is_empty());
}
}
+107
View File
@@ -0,0 +1,107 @@
//! TRACES: FR-MRG-4
//! MI-GAN, the border filler, under the inference engine.
//!
//! Sargsyan et al., ICCV 2023 (Picsart AI Research): inpainting built for
//! phones — about six million parameters of plain convolutions, no FFT and
//! no attention, so it quantises and runs on a DSP. MIT, code and weights
//! (`models/LICENCE.md`). The bare 512 generator is what ships, exported
//! at a fixed shape by `tools/export-migan.sh`; its six operator types load
//! on every rung, and what they cost is the whole story of whether a fill
//! is interactive: 7.4 s a tile under tract, 0.4 s under ONNX Runtime's
//! CPU pool, 23 ms in fp16 and 13 ms in int8 on a laptop's TensorRT
//! (2026-09-19, docs/panorama.md §12).
//!
//! The model's contract, from the reference `export_inference_model.py`:
//! input `1×4×512×512` float — channel 0 is `mask − 0.5` with 1 where the
//! picture is known, channels 1–3 the RGB in −1..1 with the unknown pixels
//! zeroed; output `1×3×512×512` in −1..1, of which the caller keeps the
//! unknown pixels. That is [`crate::fill::Inpainter`], and the rest —
//! which tiles, what context, how to blend — is `fill.rs`.
use crate::fill::Inpainter;
use crate::PanoError;
/// The tile the shipped export takes.
pub const TILE: usize = 512;
pub struct MiGan {
model: dr_inference_engine::Model,
}
impl MiGan {
/// From the model file, in whichever form the engine's rung wants
/// (`resolve_model` picks an int8 sibling for the Hexagon).
pub fn from_path(path: &std::path::Path) -> Result<Self, PanoError> {
use dr_inference_engine::{resolve_model, Role};
let (path, form) = resolve_model(Role::Inpainter, path);
let bytes = std::fs::read(&path).map_err(PanoError::ModelRead)?;
Self::from_bytes(&bytes, form)
}
pub fn from_bytes(bytes: &[u8], form: dr_inference_engine::Form) -> Result<Self, PanoError> {
use dr_inference_engine::Role;
Ok(MiGan {
model: dr_inference_engine::open(Role::Inpainter, form, bytes)?,
})
}
/// Where the fill runs, for a status line.
pub fn rung(&self) -> Result<dr_inference_engine::Rung, PanoError> {
Ok(self.model.acquire()?.rung())
}
}
impl Inpainter for MiGan {
fn tile(&self) -> usize {
TILE
}
fn fill(&mut self, rgb: &[f32], known: &[bool]) -> Result<Vec<f32>, PanoError> {
let n = TILE * TILE;
if rgb.len() != n * 3 || known.len() != n {
return Err(PanoError::Input(format!(
"MI-GAN takes a {TILE}×{TILE} tile; given {} values and {} mask entries",
rgb.len(),
known.len()
)));
}
// NCHW: the mask plane, then the three masked colour planes.
let mut input = vec![0.0f32; 4 * n];
for i in 0..n {
let m = if known[i] { 1.0 } else { 0.0 };
input[i] = m - 0.5;
for c in 0..3 {
input[(c + 1) * n + i] = (rgb[i * 3 + c] * 2.0 - 1.0) * m;
}
}
let tensor = ort::value::Tensor::from_array(
ndarray::Array::from_shape_vec(ndarray::IxDyn(&[1, 4, TILE, TILE]), input)
.expect("shape matches by construction"),
)?;
let started = std::time::Instant::now();
let acquired = self.model.acquire()?;
let acquired_at = started.elapsed();
let mut session = acquired.lock();
let outputs = session.run(ort::inputs![tensor])?;
log::trace!(
"migan: tile on {} — acquire {:.1} ms, run {:.1} ms",
acquired.rung().label(),
acquired_at.as_secs_f64() * 1e3,
(started.elapsed() - acquired_at).as_secs_f64() * 1e3
);
let (shape, data) = outputs[0].try_extract_tensor::<f32>()?;
let dims: Vec<i64> = shape.iter().copied().collect();
if dims != [1, 3, TILE as i64, TILE as i64] {
return Err(PanoError::Model(format!(
"MI-GAN output is {dims:?}, expected [1, 3, {TILE}, {TILE}]"
)));
}
let mut out = vec![0.0f32; n * 3];
for i in 0..n {
for c in 0..3 {
out[i * 3 + c] = (data[c * n + i] * 0.5 + 0.5).clamp(0.0, 1.0);
}
}
Ok(out)
}
}
+197
View File
@@ -0,0 +1,197 @@
//! TRACES: FR-MRG-4
//! The surface the composite is drawn on.
//!
//! A panorama is a set of directions; a picture is a plane. The projection
//! is the map between them, and the three offered are the three every
//! stitcher offers because each is right for a different field of view:
//! perspective keeps straight lines straight and cannot reach 180°;
//! cylindrical keeps verticals vertical and stretches nothing horizontally,
//! for the wide single row; spherical for anything that also looks up.
//!
//! Every function here is the *inverse* map — output pixel to direction —
//! because that is what a gather needs (`lens.rs` in `dr-pipeline` says
//! why a warp is written that way), and it is the function the WGSL warp
//! will repeat verbatim. The forward map exists for bounds only.
use crate::linalg::Vec3;
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Projection {
Perspective,
Cylindrical,
Spherical,
}
impl Projection {
/// Which projection a field of view calls for.
///
/// Perspective stretches the edges by `1 / cos` of the angle from the
/// centre, which is 2× at 60° and unbounded at 90°; the switch is where
/// that stretch starts to look like a mistake. Spherical is for a set
/// that spans enough vertically that a cylinder would stretch the top
/// and bottom the same way.
pub fn suggest(horizontal_fov: f64, vertical_fov: f64) -> Projection {
if horizontal_fov < 70f64.to_radians() && vertical_fov < 70f64.to_radians() {
Projection::Perspective
} else if vertical_fov < 100f64.to_radians() {
Projection::Cylindrical
} else {
Projection::Spherical
}
}
/// The direction an output point looks along. `scale` is the output's
/// focal length in pixels: the radius of the cylinder or sphere, or the
/// plane's distance. Coordinates are centred on the projection's origin
/// (the direction `+z`).
pub fn to_direction(self, scale: f64, u: f64, v: f64) -> Vec3 {
match self {
Projection::Perspective => Vec3::new(u, v, scale).normalised(),
Projection::Cylindrical => {
let theta = u / scale;
Vec3::new(theta.sin(), v / scale, theta.cos()).normalised()
}
Projection::Spherical => {
let theta = u / scale;
let phi = v / scale;
Vec3::new(theta.sin() * phi.cos(), phi.sin(), theta.cos() * phi.cos())
}
}
}
/// Where a direction lands on the output, or `None` where the
/// projection cannot show it (behind a perspective plane, at a
/// cylinder's poles).
pub fn from_direction(self, scale: f64, d: Vec3) -> Option<(f64, f64)> {
let (x, y, z) = (d.x(), d.y(), d.z());
match self {
Projection::Perspective => (z > 1e-9).then(|| (scale * x / z, scale * y / z)),
Projection::Cylindrical => {
let r = (x * x + z * z).sqrt();
(r > 1e-9).then(|| (scale * x.atan2(z), scale * y / r))
}
Projection::Spherical => {
let r = (x * x + z * z).sqrt();
Some((scale * x.atan2(z), scale * y.atan2(r)))
}
}
}
}
/// The output rectangle a set of frames covers, in centred output pixels.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct Bounds {
pub min_u: f64,
pub min_v: f64,
pub max_u: f64,
pub max_v: f64,
}
impl Bounds {
pub fn width(&self) -> f64 {
self.max_u - self.min_u
}
pub fn height(&self) -> f64 {
self.max_v - self.min_v
}
}
/// Bounds of the frames' footprints under `projection`, by walking each
/// frame's border.
///
/// `frame_size` is the frames' width and height in the same pixels the
/// cameras' focal length is in. The border is sampled rather than only its
/// corners because under a cylinder the widest point of a rolled frame is
/// not a corner.
pub fn bounds(
projection: Projection,
scale: f64,
cameras: &crate::bundle::Cameras,
frame_size: (f64, f64),
) -> Option<Bounds> {
let (w, h) = frame_size;
let mut b: Option<Bounds> = None;
let steps = 64;
for k in 0..cameras.rotations.len() {
for s in 0..steps {
let t = s as f64 / steps as f64;
for p in [
(-w / 2.0 + w * t, -h / 2.0),
(-w / 2.0 + w * t, h / 2.0),
(-w / 2.0, -h / 2.0 + h * t),
(w / 2.0, -h / 2.0 + h * t),
] {
let d = cameras.bearing(k, p);
let Some((u, v)) = projection.from_direction(scale, d) else {
continue;
};
b = Some(match b {
None => Bounds {
min_u: u,
min_v: v,
max_u: u,
max_v: v,
},
Some(b) => Bounds {
min_u: b.min_u.min(u),
min_v: b.min_v.min(v),
max_u: b.max_u.max(u),
max_v: b.max_v.max(v),
},
});
}
}
}
b
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn to_and_from_direction_are_inverses() {
for proj in [
Projection::Perspective,
Projection::Cylindrical,
Projection::Spherical,
] {
for (u, v) in [(0.0, 0.0), (300.0, -200.0), (-900.0, 450.0)] {
let d = proj.to_direction(1000.0, u, v);
let (bu, bv) = proj.from_direction(1000.0, d).expect("in front");
assert!(
(bu - u).abs() < 1e-9 && (bv - v).abs() < 1e-9,
"{proj:?} {u} {v}"
);
}
}
}
#[test]
fn the_origin_looks_down_z_in_every_projection() {
for proj in [
Projection::Perspective,
Projection::Cylindrical,
Projection::Spherical,
] {
let d = proj.to_direction(500.0, 0.0, 0.0);
assert!((d.z() - 1.0).abs() < 1e-12);
}
}
#[test]
fn a_cylinder_maps_ninety_degrees_to_a_quarter_turn_of_pixels() {
let d = Vec3::new(1.0, 0.0, 0.0);
let (u, v) = Projection::Cylindrical.from_direction(100.0, d).unwrap();
assert!((u - 100.0 * std::f64::consts::FRAC_PI_2).abs() < 1e-9);
assert_eq!(v, 0.0);
assert!(Projection::Perspective.from_direction(100.0, d).is_none());
}
#[test]
fn suggestion_widens_with_the_field() {
assert_eq!(Projection::suggest(0.5, 0.5), Projection::Perspective);
assert_eq!(Projection::suggest(2.5, 0.8), Projection::Cylindrical);
assert_eq!(Projection::suggest(3.0, 2.5), Projection::Spherical);
}
}
+151
View File
@@ -0,0 +1,151 @@
//! TRACES: FR-MRG-8
//! The XFeat detector — the network under tract, and the decoder after it.
//!
//! Apache-2.0 weights (`models/LICENCE.md`), exported at a fixed shape by
//! `tools/export-xfeat.sh` and loaded through the same `dr-inference-engine`
//! `dr-segment` and `dr-face` use, so this adds no runtime and no C to the
//! tree; what runs it is the device's business (docs/inference.md). ~300 ms
//! per frame on tract on the reference desktop, ~400 ms on the tablet
//! (S15.2, S15.4).
use crate::features::{decode_xfeat, DecodeOptions, Features, XFeatMaps, DESCRIPTOR_LEN};
use crate::image::Gray;
use crate::PanoError;
/// The two input shapes the shipped exports were made for: one landscape,
/// one portrait, the same weights. A frame is fitted into whichever
/// matches its aspect, so a portrait set does not spend half the
/// detector's width on padding — which is what the 6D fixture did before
/// the second export existed (512 × 768 of a 1024 × 768 input). A
/// different size is a different file (`tools/export-xfeat.sh`).
pub const INPUT_LANDSCAPE: (usize, usize) = (1024, 768);
pub const INPUT_PORTRAIT: (usize, usize) = (768, 1024);
/// The long edge of the detector's input, for callers sizing a proxy.
pub const INPUT_LONG_EDGE: usize = 1024;
#[cfg(feature = "embedded-model")]
const EMBEDDED_LANDSCAPE: &[u8] = include_bytes!("../../../models/keypoints/xfeat-1024.onnx");
#[cfg(feature = "embedded-model")]
const EMBEDDED_PORTRAIT: &[u8] = include_bytes!("../../../models/keypoints/xfeat-768.onnx");
/// A loaded detector: the network at both shapes.
pub struct XFeat {
landscape: dr_inference_engine::Model,
portrait: dr_inference_engine::Model,
pub options: DecodeOptions,
}
/// The bytes of both exports compiled into the binary, for whoever compiles
/// engines ahead of the first request (docs/inference.md §6).
#[cfg(feature = "embedded-model")]
pub fn embedded_model_bytes() -> [&'static [u8]; 2] {
[EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT]
}
impl XFeat {
/// The weights compiled into the binary.
#[cfg(feature = "embedded-model")]
pub fn embedded() -> Result<Self, PanoError> {
Self::from_bytes(EMBEDDED_LANDSCAPE, EMBEDDED_PORTRAIT)
}
/// From the two exports on disk.
pub fn from_paths(
landscape: &std::path::Path,
portrait: &std::path::Path,
) -> Result<Self, PanoError> {
let l = std::fs::read(landscape).map_err(PanoError::ModelRead)?;
let p = std::fs::read(portrait).map_err(PanoError::ModelRead)?;
Self::from_bytes(&l, &p)
}
pub fn from_bytes(landscape: &[u8], portrait: &[u8]) -> Result<Self, PanoError> {
use dr_inference_engine::{Form, Role};
Ok(XFeat {
landscape: dr_inference_engine::open(Role::Keypoints, Form::F32, landscape)?,
portrait: dr_inference_engine::open(Role::Keypoints, Form::F32, portrait)?,
options: DecodeOptions::default(),
})
}
/// Detect keypoints in an upright grayscale image.
///
/// The image is fitted into the network's input of matching aspect —
/// scaled down if larger, never up, and padded to the right and bottom
/// — and the keypoints come back in the coordinates of `image` itself,
/// so a caller that already scaled a frame to a proxy maps them on with
/// the scale it used and nothing else.
pub fn detect(&mut self, image: &Gray) -> Result<Features, PanoError> {
let ((in_w, in_h), model) = if image.height > image.width {
(INPUT_PORTRAIT, &self.portrait)
} else {
(INPUT_LANDSCAPE, &self.landscape)
};
let acquired = model.acquire()?;
let mut session = acquired.lock();
let (fitted, scale) = image.fitted(in_w, in_h);
let padded = fitted.padded(in_w, in_h);
let input =
ndarray::Array::from_shape_vec(ndarray::IxDyn(&[1, 1, in_h, in_w]), padded.data)
.expect("shape matches the buffer by construction");
let tensor = ort::value::Tensor::from_array(input).map_err(PanoError::Inference)?;
let outputs = session
.run(ort::inputs![tensor])
.map_err(PanoError::Inference)?;
let (w8, h8) = (in_w / 8, in_h / 8);
let expect = |i: usize, channels: usize| -> Result<Vec<f32>, PanoError> {
let (shape, data) = outputs[i]
.try_extract_tensor::<f32>()
.map_err(PanoError::Inference)?;
let dims: Vec<i64> = shape.iter().copied().collect();
if dims != [1, channels as i64, h8 as i64, w8 as i64] {
return Err(PanoError::Model(format!(
"output {i} is {dims:?}, expected [1, {channels}, {h8}, {w8}] — \
not the export this decoder was written for"
)));
}
Ok(data.to_vec())
};
let feats = expect(0, DESCRIPTOR_LEN)?;
let keypoints = expect(1, 65)?;
let heatmap = expect(2, 1)?;
let mut features = decode_xfeat(
&XFeatMaps {
feats: &feats,
keypoints: &keypoints,
heatmap: &heatmap,
width: w8,
height: h8,
},
&self.options,
);
// Back to the caller's image: drop anything the padding produced,
// undo the fit.
let border = self.options.border as f32;
let limit_x = fitted.width as f32 - border;
let limit_y = fitted.height as f32 - border;
let mut kept_kp = Vec::with_capacity(features.len());
let mut kept_desc = Vec::with_capacity(features.descriptors.len());
for (i, kp) in features.keypoints.iter().enumerate() {
if kp.x >= limit_x || kp.y >= limit_y {
continue;
}
kept_kp.push(crate::features::Keypoint {
x: (kp.x / scale as f32),
y: (kp.y / scale as f32),
score: kp.score,
});
kept_desc.extend_from_slice(features.descriptor(i));
}
features.keypoints = kept_kp;
features.descriptors = kept_desc;
features.width = image.width;
features.height = image.height;
Ok(features)
}
}
@@ -9,7 +9,6 @@
id: capture_sharpen
order: 120
attributes: [detail]
rust: CaptureSharpen
why_rust: |
-1
View File
@@ -6,7 +6,6 @@
id: clarity
order: 130
attributes: [detail]
rust: Clarity
why_rust: |
+320
View File
@@ -0,0 +1,320 @@
# TRACES: FR-DEV-12
id: colour_grading
label: op.colour_grading
order: 105
attributes: [colour]
doc: |
Colour grading — a hue and a strength for the shadows, the midtones and the
highlights, and one more cast over the whole frame.
The distinction from [`colour_mixer`](colour_mixer.yaml) is the whole reason
this exists. The mixer reaches for a hue that is *already in the picture*:
it will turn the greens that are there, and it can do nothing at all where
there are none. This reaches for a tonal *range* and casts colour into it
whether any was there or not — which makes it the only control that can warm
an already-neutral highlight, and the only one that can tone a monochrome
conversion, since a picture with no hue left in it gives the mixer nothing to
find.
Split toning is the case worth naming: cool shadows against warm highlights,
which is what a print toned in two baths did and what most of the looks sold
since are built on.
placement: |
After every colour correction, the mixer included. Grading is the closing
statement rather than a correction, so it should land on the colours the
photographer has already settled rather than be argued with by a control
further down the chain. Before the detail stage, which is where every
neighbourhood operation runs whatever this file says.
params:
shadow_hue:
label: param.shadow_hue
kind: scalar
min: 0
max: 360
default: 0
unit: none
scale: linear
precision: 0
doc: |
Degrees around the hue wheel, red at zero — the same units the straighten
control reports, and for the same reason: a photographer reading "210"
knows where on the wheel that is, where a normalised 0..1 has to be
translated first.
The two ends of the travel are the same colour. That is a fact about
hue rather than a defect, and it is what a wheel makes obvious and a
slider cannot — one of the reasons `presentation:` below asks for one.
shadow_strength:
label: param.shadow_strength
kind: scalar
min: 0
max: 100
default: 0
unit: percent
scale: linear
precision: 0
doc: |
How far from neutral, which is the wheel's radius. It does not go
negative: a negative radius is the hue opposite, which the hue control
already says, and two ways to spell one colour is how a preset comes
back looking like its own complement.
midtone_hue:
label: param.midtone_hue
kind: scalar
min: 0
max: 360
default: 0
unit: none
scale: linear
precision: 0
midtone_strength:
label: param.midtone_strength
kind: scalar
min: 0
max: 100
default: 0
unit: percent
scale: linear
precision: 0
doc: |
The midtones are where skin lives, so this is the range that goes wrong
first and the one most often left at zero. It is offered anyway because
a grade that can only reach the ends of the scale cannot answer a cast
that sits in the middle of it.
highlight_hue:
label: param.highlight_hue
kind: scalar
min: 0
max: 360
default: 0
unit: none
scale: linear
precision: 0
highlight_strength:
label: param.highlight_strength
kind: scalar
min: 0
max: 100
default: 0
unit: percent
scale: linear
precision: 0
global_hue:
label: param.global_hue
kind: scalar
min: 0
max: 360
default: 0
unit: none
scale: linear
precision: 0
global_strength:
label: param.global_strength
kind: scalar
min: 0
max: 100
default: 0
unit: percent
scale: linear
precision: 0
doc: |
The cast that carries no tonal weight — it applies equally at every
brightness. Without it, a photographer wanting one colour everywhere and
a second in one range has to set all three ranges to the first hue, and
can then no longer move any one of them without disturbing the other
two.
# The neutral is *no strength anywhere*, not "nothing has been touched".
#
# A hue with no strength behind it is a direction with no distance: the
# picture is identical, and every uniform below is zero. Under the default
# rule, nudging a hue while the strength sat at zero would make the node
# active, and it would then cost a uniform block, a helper and a block in the
# fused shader for a change nobody can see. Strengths cannot go negative, so
# their sum is zero exactly when every one of them is.
active: shadow_strength + midtone_strength + highlight_strength + global_strength
uniforms:
shadow_angle:
value: shadow_hue * 0.017453292
doc: |
Degrees to radians, here rather than in the shader. The wheel is a slider
on the photographer's side and a cosine on the GPU's, and the conversion
belongs at the seam between them — the fragment then has one meaning for
an angle rather than two.
shadow_amount:
value: shadow_strength / 100 * 0.5
doc: |
Full strength is half a stop of shift on the leading channel, the same
ceiling white balance holds itself to and for the same reason: a colour
control that can blow a channel on its own is a trap. Four of these can
stack, so the honest worst case is a stop, which takes both a maximum
global cast and a maximum range on top of it.
midtone_angle: midtone_hue * 0.017453292
midtone_amount: midtone_strength / 100 * 0.5
highlight_angle: highlight_hue * 0.017453292
highlight_amount: highlight_strength / 100 * 0.5
global_angle: global_hue * 0.017453292
global_amount: global_strength / 100 * 0.5
helpers: [luminance, tone_position]
define:
hue_cast: |
// A per-channel gain carrying a hue, centred on no change at all.
//
// The three channels are cosines 120 degrees apart, which is the hue wheel
// written directly as an RGB direction — a round trip through HSV would
// buy nothing here and would have to decide what to do with a colour that
// has no hue. Their sum is zero at every angle, so exp2 turns them into
// three gains whose product is exactly one: the cast tilts the balance
// without moving the overall level. That property is what keeps grading
// from doubling as an exposure control, which is the failure that has the
// photographer chasing brightness with a colour slider.
fn hue_cast(angle: f32, amount: f32) -> vec3<f32> {
let tilt = vec3<f32>(cos(angle), cos(angle - 2.0943951), cos(angle + 2.0943951));
return exp2(tilt * amount);
}
wgsl: |
let pos = tone_position(luminance(c));
// Three weights that partition the tonal scale: at every luminance they sum
// to exactly one. The two ends use the same 0.5 midpoint as
// highlights_shadows, so "shadows" means the same range of the picture in
// both places, and the midtones are defined as whatever the ends leave.
//
// The partition is what makes an equal setting on all three identical to
// the global cast. Weights that overlapped would make the join between two
// ranges stronger than either of them, so a split tone would darken or
// colour its own midtones as a side effect of the two settings meeting —
// an interaction with no control over it.
let lo_w = 1.0 - smoothstep(0.0, 0.5, pos);
let hi_w = smoothstep(0.5, 1.0, pos);
let mid_w = 1.0 - lo_w - hi_w;
// The ranges compose by multiplication rather than by mixing, because a
// product of luminance-neutral triples is another one — so three casts and
// a global still leave the tonal relationships the tone controls
// established. The global cast takes no weight: it is the whole frame.
c = c * hue_cast(shadow_angle, shadow_amount * lo_w)
* hue_cast(midtone_angle, midtone_amount * mid_w)
* hue_cast(highlight_angle, highlight_amount * hi_w)
* hue_cast(global_angle, global_amount);
presentation:
# Four wheels, each a hue and a distance from the centre. Named in pairs
# because that is the order a wheel wants them; a frontend that draws none
# of these renders eight ordinary sliders and the edit is unchanged, which
# is the state this ships in.
#
# That fallback is why each parameter carries its range in its own name
# instead of leaning on the wheel to say which one it belongs to. A flat
# list is what the panel produces today, and four sliders all called "Hue"
# would be four controls nobody can tell apart.
widgets: [colour_wheel]
demand:
two_dimensional: true
# A hue a few degrees off is a slightly different warm, not a wrong
# answer, so this is usable with a fingertip. Saying otherwise would take
# the wheel away from touch to protect an accuracy nobody needs from it.
precise_pointing: false
params:
- shadow_hue
- shadow_strength
- midtone_hue
- midtone_strength
- highlight_hue
- highlight_strength
- global_hue
- global_strength
tests:
- name: it_starts_neutral
why: |
Neutral means absent: at defaults the node must contribute no code, no
uniform and no branch to the fused shader. Every angle and amount being
zero is the arithmetic half of that; `expect_active` is the half that
keeps it out of the shader at all.
expect:
shadow_angle: 0.0
shadow_amount: 0.0
midtone_amount: 0.0
highlight_amount: 0.0
global_amount: 0.0
expect_active: false
- name: a_hue_with_no_strength_is_still_neutral
why: |
What `active:` buys. A hue is a direction and a strength is the distance
travelled along it, so a hue moved on its own changes nothing — but the
default rule ("some parameter has moved") would call the node active and
make every fused shader carry it for nothing.
set: { shadow_hue: 240, global_hue: 40 }
expect_active: false
- name: strength_alone_reaches_the_shader
why: |
The complement, and the reason the neutral cannot simply be "nothing
touched": red is hue zero, so a grade toward red never moves a hue
slider off its default and would otherwise never be applied.
set: { shadow_strength: 20 }
expect_active: true
- name: hue_is_converted_to_radians
why: |
The shader's cosines take radians and this uniform is the only place the
conversion happens. Degrees arriving unconverted would be an angle 57
times too large — it would wrap the wheel several times and land on a
colour with no relation to the one under the pointer, which reads as the
control being broken rather than as a missing constant.
set: { shadow_hue: 180 }
expect: { shadow_angle: 3.1415926 }
- name: full_strength_is_half_a_stop
why: |
The ceiling on one range's contribution, matching white balance's. A
grade that could take a channel to clipping by itself would make the
strength slider unusable over its top third.
set: { highlight_strength: 100 }
expect: { highlight_amount: 0.5 }
- name: strength_maps_linearly_onto_the_shift
why: |
Half the slider must be half the shift. A curve here would make the
wheel's radius mean something different at each distance from the
centre, and a wheel is read as a distance.
set: { midtone_strength: 50 }
expect: { midtone_amount: 0.25 }
- name: the_three_ranges_partition_the_tones
why: |
The midtone weight is defined as what the other two leave, so the three
sum to one at every luminance. Computed independently they would overlap
at the joins, and a split tone would then colour its own midtones as a
side effect of the shadow and highlight settings meeting.
expect_wgsl: ["let mid_w = 1.0 - lo_w - hi_w;"]
- name: the_global_cast_takes_no_tonal_weight
why: |
It is the one that means "everywhere". Weighted like the others it would
land mostly on the midtones, since that weight is the largest across the
range an ordinary photograph occupies, and "global" would quietly become
a fourth midtone control.
expect_wgsl: ["hue_cast(global_angle, global_amount)"]
- name: a_cast_does_not_change_the_level
why: |
The cosines sum to zero at every angle, so the three gains multiply to
one and a cast tilts the balance without lifting or dropping the
picture. Built any other way — a clamp, an added tint, an HSV round trip
— a strong grade would double as an exposure change, and the
photographer would correct it with a control that cannot reach it.
expect_helper_wgsl:
hue_cast: ["return exp2(tilt * amount);"]
-1
View File
@@ -1,6 +1,5 @@
id: colour_mixer
order: 100
attributes: [colour]
rust: ColourMixer
why_rust: |
+56
View File
@@ -0,0 +1,56 @@
# A hand-written node, and a neighbourhood one: the veil it removes is measured
# from the pixels around the one it is writing, so it runs in the detail stage
# rather than as a fragment in the fused pass. See `../src/detail.rs` for why
# that stage exists and `README.md`'s "Nodes that read their neighbours" for
# the contract.
#
# As with every `rust:` node, its descriptor, parameters and behaviour come
# from the type; this file exists so that `ops/` remains the one place the
# pipeline's order is written down.
id: dehaze
order: 125
rust: Dehaze
why_rust: |
Haze is defined by what the pixels around a pixel are doing, and the schema
above describes a function of one colour — `wgsl:` is handed `c` and no
coordinate, which is the wall the detail stage exists on the other side of.
It declares `Affects::Detail` and returns five `DetailPass`es: four that
erode the dark channel into a per-pixel veil, and one that inverts the
scattering model with the transmission that veil implies.
Nor is it four facts. The erosion is split into two exact stages per axis so
that a patch 1% of the frame wide costs its square root in taps, and the
split is arithmetic over the render size that has to be recomputed every
frame. Stretching this schema to express it would produce a worse language
than Rust, aimed at one caller.
placement: |
First among the compositional detail nodes: after noise reduction and
capture sharpening, before clarity and texture.
After noise reduction because dehaze divides by a transmission below one, so
it amplifies whatever noise is in the veiled distance by exactly the factor
it recovers the contrast by. Running it first would ask the denoiser to
remove grain that dehaze had already multiplied — the same argument clarity
records, and stronger here, because the amplification is largest in the low-
contrast regions where noise is most visible.
Before clarity and texture, and for the reason that orders those two against
each other: coarse before fine. Dehaze acts on the widest structure in the
frame, the veil that varies with distance, and clarity's base should be
computed on the picture as the veil has left it rather than on a modelling
that is about to be divided out.
What this cannot honour, and it is worth writing down rather than leaving to
be rediscovered: dehaze shifts colour. It subtracts a grey term and rescales,
so it changes saturation everywhere the veil is thick, and the colour
controls would ideally be correcting the picture that leaves here. They
cannot be. The detail stage runs as a *group* after every point operation,
because a neighbourhood pass is a separate dispatch reading a texture the
fused pass has already finished writing — so an `order:` placing this node
ahead of `vibrance` or `colour_mixer` would be a lie the chain cannot tell.
Interleaving the two would mean splitting the fused pass in half around this
one, which costs a second full-frame dispatch and intermediate on every edit
in the catalogue, whether or not it uses dehaze at all.
+4 -1
View File
@@ -1,6 +1,9 @@
id: film_sim
order: 25
attributes: [tone, colour]
# What this node is *about* is not written here, and cannot be: a `rust:` node
# publishes its own descriptor, so `attributes:` in this file would be read,
# validated and then ignored. See `Attribute::Effect` on `FilmSim`'s descriptor
# in `../src/ops/film_sim.rs`.
rust: FilmSim
why_rust: |
@@ -8,7 +8,6 @@
id: noise_reduction
order: 110
attributes: [detail]
rust: NoiseReduction
why_rust: |
-1
View File
@@ -2,7 +2,6 @@
id: texture
order: 140
attributes: [detail]
rust: Texture
why_rust: |
-1
View File
@@ -12,7 +12,6 @@ order: 60
# Both, and this is the case the plural exists for: the RGB curve is
# tonal and the per-channel curves are chromatic. Filing it under one
# would hide it from half the people looking for it.
attributes: [tone, colour]
rust: ToneCurve
why_rust: |
+39
View File
@@ -0,0 +1,39 @@
# A hand-written node, and the first thing in `Attribute::Optics`.
#
# Written, tested and unreferenced until now: `src/ops/vignetting.rs` has
# existed with a full descriptor and a working polynomial, and without an entry
# here it was never in `chain()` — so it reached no photograph and no panel.
# This file is the whole of what was missing, which is the point of `ops/`
# being the one place the pipeline's order is written down.
id: vignetting
order: 5
rust: Vignetting
why_rust: |
It carries a lens profile's `pa` coefficients, which are not parameters: they
come from the body and lens that took the photograph, not from the
photographer, and no `uniforms:` expression could produce them. The one
parameter that *is* theirs — the manual trim — is composed with `k1` in Rust,
because the profile and the trim have to reach the shader as a single
polynomial rather than as two the fragment would have to add up.
placement: |
First in the chain, ahead of white balance and exposure.
Vignetting is what the lens did to the light before the sensor measured it,
so undoing it belongs with reading the file rather than with editing the
picture — everything downstream is then working on the frame the lens would
have delivered had it been even.
The ordering is load-bearing rather than tidy. Correcting a corner means
*dividing* by an attenuation below one, which pushes those pixels up: a fast
prime wide open needs about two stops there. Run after the tonal stages, that
recovery happens once the highlights have already been rolled off and
clipped, so it lifts values that no longer have anywhere to go and the
corners posterise instead of brightening. Run here, the headroom to hold them
still exists (ARCH §5.2).
Before distortion and chromatic aberration in intent, though those are
`lens::Warp`s rather than nodes and compose ahead of the fetch, so no `order:`
relates the two.
+27
View File
@@ -32,6 +32,33 @@ params:
label: param.tint
kind: amount
# An eyedropper, where the frontend has a canvas to hang one on.
#
# Sampling a neutral is the first move of the global tonal pass — every colour
# judgement afterwards is measured against where the grey was put — and it is
# a thing you do by pointing at the photograph, not by guessing at two sliders
# until a wall stops looking green.
#
# A *hint*, on the usual terms: the two parameters below stay ordinary
# addressable scalars, and a frontend with nowhere to host a sampler renders
# them as the sliders they already were. This one is additive rather than a
# replacement — the picker writes temperature and tint and the photographer
# still nudges them afterwards — which is a difference the frontend draws for
# itself; nothing here has to say it.
#
# The order matters and is the widget kind's own contract: the first parameter
# trades red against blue, the second green against magenta.
presentation:
widgets: [white_point]
params: [temperature, tint]
demand:
# A neutral is a point on the picture, so both axes at once.
two_dimensional: true
# Deliberately false. A grey card, a cloud, a white wall — the things
# worth sampling are large, and FR-UI-7 grows the hit region to the
# modality in any case, so a thumb is as workable as a mouse.
precise_pointing: false
# Temperature trades red against blue; tint trades green against magenta.
# Both are scaled so the full range is a strong but not destructive
# correction: ±0.5 in log2 at the extremes — half a stop of channel shift,

Some files were not shown because too many files have changed in this diff Show More