The engine knew f32 and int8, and gave the Hexagon int8 for every role it served. Measured on the tablet itself (inference.md §1.5), int8 lost 5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px, emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now names one per role: detectors and landmarks A16W8, the segmenter, scene model, border filler and denoiser A16W16, XFeat int8. The embedder and the eye classifiers stay on the CPU. Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and XFeat, compiled into the binary, embed their quantised forms on Android only and pick through `choose_embedded`. The probe, the compile step and the cache fingerprint follow the form instead of assuming int8. Detectors on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers for all three spellings. On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the same inputs, and against the CPU's f32 time: SCRFD 500m/2.5g/10g A16W8 100% of faces in every band 4.2/5.1/9.0 ms vs 17/56/198 landmarks A16W8 0.25 px in the 192 crop 0.5 ms vs 2.8 YOLO26n-seg A16W16 98.2% found, mask IoU 0.994 12.9 ms vs 90 scene model A16W16 98.9% of cells agree 15 ms vs 151 MI-GAN A16W16 41 dB from f32 in the fill 87 ms vs 488 XFeat int8 pano alignment 0.45 px (f32's own spread 0.41) 6.5 ms vs 58 denoiser A16W16 0.00 dB at every ISO 95 ms vs 1510 a tile Face numbers are over public COCO val2017 photographs, not a library. The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors removed), about 43 MB more. The Windows installer and its CI count skip them; the Arch and Flatpak packages list their files and never had them. The ladder example takes a role per model, which is how the per-role forms above were seen landing on the NPU from the real probe.
699 lines
42 KiB
Markdown
699 lines
42 KiB
Markdown
# Panorama
|
||
|
||
**Status:** Draft · 2026-09-19
|
||
**Companion to:** [requirements.md](requirements.md) §3.11 FR-MRG-1 … 11, D18, S15 · [architecture.md](architecture.md) §5.2, §6.2
|
||
|
||
The first merge (§3.11): several frames, rotated about one point, become one
|
||
photograph. This document is how that lands on the pipeline that exists now —
|
||
which stages, where each runs, how the composite is produced in chunks when it
|
||
is larger than any texture or any memory, what is ported from where, and what
|
||
the keypoint model may be under D8.
|
||
|
||
---
|
||
|
||
## 1. Why it is worth the work
|
||
|
||
The audience shoots panoramas and leaves the application to stitch them. That
|
||
is the same workflow break dust was (`spot-removal.md` §1): a RAW editor that
|
||
does everything but the one thing, and the photographer's work ends up in a
|
||
JPEG produced by a tool that never saw the RAW.
|
||
|
||
It is also the merge whose alignment problem is smallest. A panorama is a
|
||
rotation — three parameters per frame plus a focal length — with no depth to
|
||
recover. HDR merge and focus stacking share its data model (D18) and most of
|
||
its machinery (FR-MRG-3, 5, 6, 7, 10, 11 are written to be general); building
|
||
the panorama first builds the shared part on the easiest geometry.
|
||
|
||
## 2. Non-goals
|
||
|
||
- **Not structure-from-motion.** No translation is solved for. A hand-held set
|
||
with parallax gets its ghosts hidden by seam placement, and a set with real
|
||
parallax is not a panorama. COLMAP's front end is the right mental model;
|
||
its back end is the wrong problem.
|
||
- **Not a multi-source Version.** D18. The composite is a file, and nothing in
|
||
the catalog, the sidecar format or sync learns about cross-references.
|
||
- **Not boundary fill.** Painting pixels that were never captured is the pixel
|
||
editing §1.3 excludes. Auto-crop is the tool.
|
||
- **Not automatic.** The tool proposes an alignment and writes nothing until
|
||
the photographer confirms. Same rule as spot removal and D17, for the same
|
||
reason: a merge that silently omits or misplaces a frame is the failure this
|
||
application must not have.
|
||
- **Not HDR-panorama in one pass.** Until HDR merge exists on its own, a
|
||
bracketed panorama is bracketed frames merged first, then stitched.
|
||
|
||
## 3. What is new, precisely
|
||
|
||
Nearly all of it, unlike spot removal. The pipeline renders one source to one
|
||
texture; nothing in the tree detects keypoints, estimates a rotation, warps
|
||
into a projection, finds a seam, or blends a pyramid. What exists and is
|
||
reused:
|
||
|
||
| Exists | Where | Reused for |
|
||
|---|---|---|
|
||
| Render a source through the fused pass, with a linear f16 output mode | `dr-gpu` demosaic → `AdjustPass`, `OutputMode::LinearWorking` | FR-MRG-2's camera-space input, as a compose entry with no operations and the profile uniforms neutral (S15.3) |
|
||
| Tiled rendering with a priority scheduler | ARCH §5.3 | Pulling source tiles on demand into an output chunk (§5 below) |
|
||
| A non-CFA source entering the pipeline | `Demosaicer::from_rgba8` | The composite's decode path, if the container is a TIFF (S15.1) |
|
||
| DNG matrices read through rawler | `dr-decode::profile` | The composite's decode path, if the container is a DNG |
|
||
| Static-shape ONNX under tract, heads decoded in Rust | `dr-segment` | The keypoint model (§6) |
|
||
| A batch worker with its own `GpuContext`, activity row, cancel | `dr-ui::export` | FR-MRG-7 verbatim |
|
||
| The 16-bit TIFF encoder with metadata sub-IFDs | `dr-export::encode` | FR-MRG-3's writer, extended to linear samples |
|
||
| Multi-select in the grid | `collections_ui::selected` | The entry point |
|
||
|
||
New: a `core/dr-pano` crate holding the geometry (keypoints, matching, the
|
||
rotation solve), a set of WGSL passes in `dr-gpu` (reprojection, gain,
|
||
seam, pyramid blend), the chunked output driver, the container writer, and
|
||
the dialog.
|
||
|
||
## 4. The stages, and where each runs
|
||
|
||
FR-MRG-10 states the rule; this is the table it was written from.
|
||
|
||
| Stage | Cost shape | Runs on | Why |
|
||
|---|---|---|---|
|
||
| Source to camera-linear | per pixel, full res | GPU, the existing pipeline | It *is* the pipeline, stopped early |
|
||
| Keypoint detection | once per frame, at 1024 px | CPU, tract (NEON on the tablet) | Bounded by frame count, not output size. Same runtime faces and masks use. Hand-written WGSL convolutions for a model that runs five times would be work with no visible gain. |
|
||
| Descriptor matching | K² × D per pair | CPU, SIMD | 2048² × 64 × 10 pairs ≈ 3 GFLOP — tens of milliseconds |
|
||
| Rotation solve, bundle adjustment | 3N + 1 parameters, Levenberg–Marquardt | CPU | Microseconds. Not parallel work. |
|
||
| Preview reprojection | per pixel, proxy res | GPU, interactive | Projection and horizon changes re-warp N proxies at frame rate |
|
||
| Full-resolution warp | per output pixel | GPU, chunked (§5) | The heaviest thing in the application |
|
||
| Gain compensation | per overlap region | GPU reduction, then N scalars | Sums, on the histogram pass's pattern (ARCH §5.5) |
|
||
| Seam finding | per overlap pixel | GPU-friendly variant | Graph cut resists the GPU; a distance-transform or per-column DP seam does not. The algorithm is chosen for the GPU, not for the paper. |
|
||
| Multi-band blend | per pixel × levels | GPU, chunked | Laplacian pyramids are separable convolutions — the detail stage's shape |
|
||
| Encode | per pixel, once | CPU, streamed per chunk row | As export does |
|
||
|
||
## 5. Chunked in output space
|
||
|
||
FR-MRG-11 forbids holding the composite as one texture, and two facts force it
|
||
before memory does:
|
||
|
||
- `max_texture_dimension_2d` is 8192 on many mobile GPUs and 16384 on desktop.
|
||
A three-row panorama is routinely 20 000 px wide.
|
||
- Five 24 MP frames at working precision are ~1 GB together. The tablet does
|
||
not have it.
|
||
|
||
**The geometry is known before any full-resolution pixel exists.** Alignment
|
||
runs on proxies; what comes out is a rotation per frame, a focal length, a
|
||
projection and an output rectangle. From those, every output pixel's source
|
||
coordinates in every frame are a closed-form function. That is what makes
|
||
chunking simple rather than clever:
|
||
|
||
```
|
||
for each output chunk C (e.g. 2048 × 2048, in output space):
|
||
frames_in(C) = frames whose projected footprint intersects C
|
||
for each frame F in frames_in(C):
|
||
source tiles T(F, C) = tiles of F that project into C, plus a margin
|
||
render T(F, C) to scene-linear through the pipeline's tile cache
|
||
warp T(F, C) into C's coordinate frame ← GPU
|
||
gain-correct, seam, blend within C ← GPU, with overlap
|
||
read C back, encode its rows ← CPU, streamed
|
||
```
|
||
|
||
The working set is one chunk, its per-frame warped copies, and the source
|
||
tiles that fed them. It does not grow with the composite.
|
||
|
||
**The blend needs a margin.** A Laplacian pyramid of L levels reads
|
||
2^L pixels beyond the chunk edge; a chunk is therefore rendered with a margin
|
||
of that width and the margin discarded after the blend. Seams cross chunk
|
||
boundaries and must agree on both sides: the seam is found once at a reduced
|
||
resolution over the whole overlap (which fits — it is a mask, not an image),
|
||
then upsampled into each chunk. The same is true of gain: the scalars are
|
||
solved once from proxy-resolution overlaps and applied everywhere.
|
||
|
||
**Source tiles are the pipeline's tiles.** ARCH §5.3's cache keys by
|
||
`(VersionId, tile, zoom, graph_hash_prefix)`; the merge asks for tiles of a
|
||
neutral graph at zoom 1 and gets the same caching every other consumer does.
|
||
A tile pulled for one chunk is usually needed by the neighbouring chunk, and
|
||
stays hot for it.
|
||
|
||
### 5.1 The tap — S15.3, answered by reading the composer
|
||
|
||
The fused shader's order, fixed by `operation.rs`'s own tests: warp → as-shot
|
||
white balance → operations → base curve → camera matrix → store. *(Amended
|
||
2026-09-27, D19: warp → as-shot white balance → white balance → camera matrix
|
||
→ operations → view transform → store. The tap is unaffected: it has no
|
||
operations, its caller fills the matrix with the identity, and the composer
|
||
emits no view transform in `OutputMode::CameraLinear`.)* The store is
|
||
either the display encode or, in `OutputMode::LinearWorking`, an unclipped
|
||
`rgba16float` of linear sRGB. That mode exists for the detail stage and is
|
||
selected from the operations, never by a caller flag, so that a shader and
|
||
the texture bound to it cannot disagree.
|
||
|
||
The merge wants the values *before* the curve and matrix (FR-MRG-2), and the
|
||
composer already makes that a matter of uniforms rather than structure: the
|
||
white balance, the matrix and the curve's active flag are all in the reserved
|
||
uniform block, and a fused pass with no operations, `as_shot_wb = 1`,
|
||
`cam_to_srgb = I` and `base_curve_last.z = 0` stores exactly camera-linear
|
||
RGB after the warp. *(Since D19 there is no curve flag: the base curve is
|
||
gone, and the tap composes no view transform, so the white balance and the
|
||
matrix are the only uniforms it fills neutral.)* So the tap is:
|
||
|
||
- `EditGraph::compose_camera_linear()` — the `LinearWorking` tail with an
|
||
empty operation list and identity framing, paired by name with
|
||
- `AdjustPass::render_camera_linear()` — binds the f16 target, fills the
|
||
reserved uniforms neutral instead of from the source, returns the texture,
|
||
- and a float readback beside the existing 8-bit one.
|
||
|
||
Nothing in the chain moves. **Precision:** the tap and every chunk buffer
|
||
after it should be `rgba32float`, not f16. A 14-bit sensor has 16 384 steps
|
||
to white; f16 has 2 048 in the top octave, and a composite that is going to
|
||
be re-developed deserves the sensor's precision. The cost is 2× on buffers
|
||
FR-MRG-11 already bounds.
|
||
|
||
**What the DNG carries as a consequence:** the first source's `Make`,
|
||
`Model` and `UniqueCameraModel` — so `base_curve::for_body` found the 6D's
|
||
curve, until D19 retired the per-body curves — its `ColorMatrix1`/`2` with illuminants, and its `AsShotNeutral`. The
|
||
composite then develops through the same profile as its sources, applied
|
||
once. The spike's 64 × 48 file (§8) already carries the matrix and neutral;
|
||
the body name is a string.
|
||
|
||
## 6. The keypoint model
|
||
|
||
FR-MRG-8: works without weights, better with them. The licence read comes
|
||
first (D13's lesson, S15.2).
|
||
|
||
| Model | Licence | Fits tract? | Position |
|
||
|---|---|---|---|
|
||
| **XFeat** (CVPR 2024) | Apache-2.0 | Plain convolutions, fully convolutional, the repo ships an ONNX export | **Chosen.** Fixed 1024 px input, dense heatmap and descriptor map out, NMS and top-K in Rust — the yolo26 pattern |
|
||
| DISK | Apache-2.0 | U-Net, static | Second choice; stronger descriptors, ~3–4× the compute |
|
||
| ALIKE | BSD-3 | Plain convolutions | Fallback if XFeat's export fails F6 |
|
||
| ALIKED | BSD-3 | Deformable convolution in the descriptor head | Unlikely to load |
|
||
| SuperPoint, SuperGlue, R2D2, SiLK, MASt3R | non-commercial | — | Out on licence |
|
||
| LightGlue | Apache-2.0 | Transformer over a variable keypoint count | Not until mutual-nearest-neighbour matching fails on a real set |
|
||
|
||
**S15.2, 2026-09-19: XFeat loads under tract.** `tools/export-xfeat.sh`
|
||
exports the network alone at 768×1024 — thirteen operator types, all
|
||
standard: `Conv`, `InstanceNormalization`, `AveragePool`, `Resize`, `Slice`,
|
||
`Transpose`, `Reshape`, `Concat`, `Add`, `Relu`, `Sigmoid`, `ReduceMean`,
|
||
`Unsqueeze` — and
|
||
[`examples/onnx_probe.rs`](../../core/dr-segment/examples/onnx_probe.rs) loads
|
||
the 2.8 MB file through the app's own `ort`-over-tract backend with nothing
|
||
unsupported, in 28 ms, and runs it in **~300 ms on the reference desktop's
|
||
CPU**. The weights ship as `models/keypoints/xfeat-1024.onnx`, recorded in
|
||
`models/LICENCE.md`. **On the tablet** (S15.4's CPU half, same day):
|
||
`tools/onnx-probe-on-device.sh` cross-builds the probe, and the same file
|
||
runs in **~400 ms per frame** on the reference tablet's NEON cores (ROD2-W09,
|
||
SM8635), with output ranges identical to the desktop's — inside NFR-MRG-1's
|
||
1 s per frame with room to spare, and 1.3× the desktop rather than the 2×
|
||
faces.md §9 measured for its scan. Still to do: a keypoint-level comparison
|
||
against the PyTorch reference once the Rust decoder exists — the probe
|
||
proves the graph runs, not that the numbers match.
|
||
|
||
The outputs are three maps at 1/8 resolution, 96×128 for the export size:
|
||
64-channel descriptors, 65-channel keypoint logits (each 8×8 cell's position
|
||
plus "none"), and a reliability heatmap. The Rust decoder is: softmax over the
|
||
65, pixel-shuffle the first 64 to full resolution, 5×5 non-maximum
|
||
suppression, top-k by reliability, bilinear sampling of the descriptor at
|
||
each keypoint, L2 normalise. That is `detectAndCompute` in the reference,
|
||
minus the network.
|
||
|
||
Without weights: AKAZE (BSD, `akaze` from rust-cv), which is adequate on
|
||
well-textured overlaps and worse on sky, repeated structure and exposure
|
||
drift — which is where a learned detector earns its place.
|
||
|
||
Matching is mutual nearest neighbour with a ratio test, then RANSAC on a
|
||
rotation model. For a panorama — one lens, near-pure rotation, 20–40 %
|
||
overlap — that is what Hugin and OpenCV's stitcher use, and it is enough.
|
||
|
||
## 7. What is ported from where
|
||
|
||
Nothing is linked; everything is read.
|
||
|
||
| Source | Licence | Taken |
|
||
|---|---|---|
|
||
| OpenCV `modules/stitching` | Apache-2.0 | The stage layout — Brown & Lowe (2007) as a set of small classes with one job each — and the warpers' projection maths |
|
||
| OpenPano (ppwwyyxx) | MIT (verify on read) | The estimation and bundle-adjustment maths, function by function, with outputs diffed against it |
|
||
| enblend-enfuse | GPLv2+ | Seam-line optimisation and Burt–Adelson multi-band blending |
|
||
| Hugin `nona` | GPLv2+ | The GLSL remapper, as the reference for the WGSL warp |
|
||
|
||
The golden set (§8 of the requirements) is OpenCV's stitcher on the same
|
||
inputs: a reference output to compare against, within a tolerance calibrated
|
||
the way S9 calibrates R1.
|
||
|
||
## 8. The output file
|
||
|
||
FR-MRG-3. A linear DNG at the source's native scale: `u16` samples on the
|
||
first source's black-subtracted scale, `WhiteLevel` = its white minus its
|
||
black (13 023 for the 6D set: 15 070 − 2 047), `BlackLevel` = 0. Not rescaled
|
||
to 65 535 — the sensor had 14 bits and the file says so, and a value the
|
||
sensor could not have produced is not invented by a multiply. The first
|
||
source's `Make`, `Model`, `UniqueCameraModel`, `ColorMatrix1/2`,
|
||
`CalibrationIlluminant1/2`, `AsShotNeutral` and EXIF are carried, so the
|
||
composite develops through the same profile as its sources. Named from the
|
||
first source with a `-pano` suffix, beside it.
|
||
|
||
Three samples per pixel rather than a CFA: the warp resamples, and there is no
|
||
sensor grid to mosaic back onto. Nothing else about being a RAW is lost —
|
||
no white balance, no curve, no matrix, no clip has been applied — and the
|
||
photographer develops the panorama afterwards as one photograph. One sample
|
||
is rewritten: a blown one, which is written as the camera value the
|
||
composite's balance calls grey rather than as the sensor's (1, 1, 1) — §15.
|
||
|
||
The sources are portrait frames in the 6D set: `Orientation` is applied
|
||
before alignment (learned features are not rotation-invariant) and the
|
||
composite is written upright with `Orientation = 1`.
|
||
|
||
Two containers were candidates and S15.1 decided, on 2026-09-19:
|
||
|
||
- **Linear DNG.** `PhotometricInterpretation = LinearRaw`, three samples per
|
||
pixel, `ColorMatrix1` carried from the first source. Re-enters through
|
||
rawler as `Format::Dng` with no new decode path, *if* rawler reads it back.
|
||
What Lightroom writes.
|
||
- **Float TIFF.** `SampleFormat = IEEEFP`, 16 or 32 bits, an ICC profile for
|
||
the working space. Needs `Format::Tiff` and a decode path, but the writer is
|
||
the existing encoder with a different sample type, and nothing about it is
|
||
uncertain.
|
||
|
||
**Linear DNG.** [`examples/linear_dng.rs`](../../core/dr-decode/examples/linear_dng.rs)
|
||
hand-rolls a 64 × 48 `LinearRaw` DNG — one IFD, uncompressed 16-bit RGB,
|
||
`DNGVersion`, `ColorMatrix1`, `AsShotNeutral`, `CalibrationIlluminant1` — and
|
||
rawler 0.7 reads it back: `cpp 3`, the samples interleaved as written, the
|
||
matrix parsed into the camera definition, and `CameraProfile::extract` builds
|
||
the same profile it would for a camera file. ImageMagick's libraw reads the
|
||
same bytes. What does *not* yet work is `dr_decode::decode`, which accepts the
|
||
file as CFA and hands the pipeline three times the samples it expects: the
|
||
`cpp == 3` branch is the work, and it is the only decode work.
|
||
|
||
The composite therefore enters the pipeline as a non-CFA, *linear* source —
|
||
`from_rgba8`'s sibling with `non_linear = false` and the colour matrix carried
|
||
from the DNG — and is developed as any RAW is. The writer is the example's
|
||
IFD, grown up: tiled rather than one strip (FR-MRG-11 encodes per chunk), and
|
||
carrying the first source's EXIF in a sub-IFD as `dr-export` already does.
|
||
|
||
## 9. Interaction
|
||
|
||
- The entry is the grid's selection: two or more images, one action, "Merge
|
||
to panorama". One image, or images from different roots, and the action
|
||
says why it is unavailable.
|
||
- The dialog shows the aligned proxies in the chosen projection, with the
|
||
projection, horizon and crop controls of FR-MRG-4, and the per-frame
|
||
residuals. A frame that failed to align is named there (FR-MRG-5), and the
|
||
merge cannot be confirmed with it in the set. *(Since 2026-09-27 each row
|
||
has a box: an unticked frame is left out and the rest are solved again from
|
||
what the first pass measured — §15.)*
|
||
- Confirm starts the FR-MRG-7 job. The composite appears in the grid when the
|
||
file is written and catalogued, beside its sources, with the merge as the
|
||
first entry in its history.
|
||
|
||
**How it gets there (FR-MRG-6, 2026-09-28).** A rescan fired as the merge
|
||
finished raced the upload it followed — the 800 MB copy into a folder library
|
||
was still running when the folder was listed, and a Nextcloud upload takes
|
||
minutes — so the listing lacked the composite, recorded the folder's
|
||
validator, and the grid did not show it until the next sync pass. Now:
|
||
|
||
- *Catalogued by the merge.* `MergeEvent::Done` carries a `Composite` — the
|
||
name it will have, the size of the picture it opens on (the crop, or the
|
||
whole when filled), the capture time written into the DNG (the mean of the
|
||
frames'; the sources' earliest where none has one), the body, and its
|
||
thumbnails. `library::catalogue_composite` writes the row in one
|
||
transaction, keyed on `(root_id, source_ref)` exactly as the scan will list
|
||
the file, at `metadata_state = 2`, and the grid reloads. The name is chosen
|
||
against the catalog's names in that folder (`names_in_folder`), since the
|
||
upload replaces whatever is at its name.
|
||
- *The server's half after the upload.* Once a file the catalog already has
|
||
a row for is sent, the drain lists its folder once, records the file id the
|
||
server assigned (`record_uploaded`) and puts the merge's thumbnails in the
|
||
store under it; then the grid rescans. A scan that ran before the upload
|
||
leaves the row alone, and the one after it updates it in place.
|
||
- *Thumbnails from the merge.* The bands are box-reduced as they are written,
|
||
after the fill, to a copy 4096 pixels long (`merge_thumbs::Reduced`). That
|
||
copy is written as a linear DNG in memory with the composite's own profile,
|
||
header and crop and opened through `open_session` — develop's first open:
|
||
the D19 pipeline, the default view transform and tone mapping, the as-shot
|
||
balance and the working-space-to-display conversion. The grid, large and
|
||
wide classes are rendered from that session, staged in the outbox as
|
||
`x.dng.thumbs` before the rename releases the payload, and drawn from
|
||
memory until the upload has a file id to store them under. A test develops
|
||
a synthetic composite both ways and holds the mean, 95th and 99.5th luma
|
||
percentiles within 3–4 levels; the naive balanced-and-gamma picture misses
|
||
by 13. Older composites, which have no staged thumbnails, are thumbnailed
|
||
the ordinary way.
|
||
- *A wide cell.* `library_ui::layout` places the grid as a lattice of slots.
|
||
`natural_span` maps aspect to 2, 3 or 4 columns (from 1.9, 2.45 and 3.46 —
|
||
√(s(s+1)) is where two neighbouring classes leave the same share of their
|
||
cell empty), capped at the columns there are and the whole row on the
|
||
tablet, and the same number names the thumbnail class (`Wide2`–`Wide4`, 512
|
||
pixels of long edge per column). A wide cell that does not fit in the rest
|
||
of a row starts the next; nothing later moves into the gap, so ordinals —
|
||
the arrows, a shift-click's run, the timeline, burst folding — are
|
||
untouched, and up/down step by rows through the layout. The window's own
|
||
read carries `w` and `h`; where the wide ones sit in the whole list is one
|
||
query, run when the list changes, and a library with no panorama answers it
|
||
from the partial index `images_wide`, created on first use rather than by a
|
||
schema bump.
|
||
|
||
## 10. Order of work
|
||
|
||
1. **S15**, all four, before anything else. (1) and (2) are a day each and
|
||
either can change the design.
|
||
2. `dr-pano`: keypoints (AKAZE first, XFeat when S15.2 passes), matching,
|
||
RANSAC, rotation solve. Unit-tested against synthetic rotations of one
|
||
frame, where the answer is known exactly.
|
||
3. The working-space tap, and the preview reprojection pass. At this point the
|
||
dialog can show an alignment.
|
||
4. The chunked driver with a feathered blend — the whole path end to end,
|
||
writing a file, before the blend is good.
|
||
5. Gain, seams, multi-band.
|
||
6. The container, the catalog entry, provenance, the history entry.
|
||
7. Tablet: NFR-MRG-1's figure, and FR-MRG-9's ceiling.
|
||
|
||
## 11. Where it stands — 2026-09-19, end of the first day
|
||
|
||
Built, on branch `merge/panorama`, in the order §10 gave:
|
||
|
||
| Piece | Where | State |
|
||
|---|---|---|
|
||
| Geometry: keypoints, matching, homography, focal, bundle adjustment, projections | `core/dr-pano` | Done; 33 tests without a model; the fixture aligns in 4.5 s |
|
||
| XFeat at two shapes under tract | `models/keypoints`, `dr_pano::xfeat` | Done; 300 ms/frame desktop, 400 ms tablet |
|
||
| The camera-space tap | `OutputMode::CameraLinear`, `AdjustPass::render_camera_linear` | Done, `rgba32float`, tiles by view rect |
|
||
| Linear DNG writer, streamed | `dr_export::write_linear_dng` | Done; rawler reads it back |
|
||
| A three-sample `RawImage` re-entering the pipeline | `dr-decode`, `DemosaicedImage::from_linear_rgb16` | Done |
|
||
| Warp, accumulate, resolve, chunk by chunk | `dr_gpu::MergePass`, `merge.wgsl` | Done; seams (§11.1) over a feather, scalar gain |
|
||
| The job: load, proxies, align, gains, confirm, merge, provenance | `dr_ui::merge` | Done; `examples/merge.rs` drives it headless |
|
||
| The page: table, preview, projection, Merge/Stop/Back; the grid's button | `merge.slint`, `merge_ui.rs` | Done; `DARKROOM_START_MERGE=a.CR2,b.CR2` lands on it |
|
||
| Placement beside the sources through the outbox, rescan | `merge_ui.rs` | Done, untested against a server |
|
||
|
||
**Measured on the fixture (desktop, 12 × 20 MP, Intel adapter):** proxies
|
||
and keypoints 4 s, alignment 4.5–12.6 s (load-sensitive: the matcher is
|
||
every core), the merge **26 s for a 22 993 × 5 980 composite** in twelve
|
||
bands of 2048 × 512 chunks, 45 s all told, an 825 MB DNG. NFR-MRG-1's 60 s
|
||
holds on the desktop with room; the tablet's figure is still S15.4's open
|
||
half.
|
||
|
||
**Open, in the order they matter:**
|
||
|
||
1. **Auto-crop (FR-MRG-4).** The merge returns a coverage mask per band and
|
||
the file carries the black border. The largest inscribed rectangle over
|
||
the coverage, then the DNG's `DefaultCropOrigin`/`DefaultCropSize`, so
|
||
nothing is thrown away and the develop view opens on the picture.
|
||
2. **The pyramid** (§10 step 5). Seams landed 2026-09-30 (§11.1); the
|
||
blend across them is one width for every frequency, so an exposure step
|
||
the gains leave is narrowed to the seam's 64 px rather than hidden over
|
||
the old 200. A Laplacian pyramid would blend low frequencies wide and
|
||
detail narrow.
|
||
3. **Vignetting in the tap.** The lens profile's distortion is applied
|
||
before the fetch; its vignetting is an operation and is not. Frame edges
|
||
are darker than their centres by the lens's falloff, and the feather
|
||
averages them into the overlaps.
|
||
4. **The tablet:** memory (twelve 40 MB sensor buffers on the CPU, one
|
||
demosaiced frame at a time on the GPU), the figure, and FR-MRG-9's
|
||
ceiling.
|
||
5. **Horizon and drag-to-correct (FR-MRG-4, the proposed 4a).** The
|
||
alignment failed on nothing in the fixture; the interaction waits for a
|
||
set it fails on.
|
||
6. **`derived_from` names sources by file name**, not content hash: the
|
||
catalog's `content_hash` is null for most images most of the time. The
|
||
hash can join it when the catalog has one.
|
||
|
||
### 11.1 Seams — 2026-09-30
|
||
|
||
The feather averaged every overlap over 200 px, so anything the frames
|
||
disagreed on — parallax on the near slope, a walker, wind in a branch — came
|
||
out twice at half strength: a soft double edge at 1:1, reported as a glitch.
|
||
|
||
`dr_pano::seam` now chooses, per output texel at proxy resolution, which
|
||
frame it is taken from. Frames are laid down nearest-first; where a new one
|
||
overlaps the composite, each texel costs the gain-corrected difference
|
||
between the two, plus the detail either has there, plus nearness to either
|
||
frame's edge (vignetting, the lens correction's fringe), taken as the
|
||
**worst** over a 4-texel window so the path stays a blend radius clear of a
|
||
difference rather than grazing it. The cut is a dynamic-programming path
|
||
across the overlap, perpendicular to the line from the composite's frames to
|
||
the new one: §4's per-column seam, not a graph cut. The map is computed per
|
||
projection, for the page's preview and again for the merge.
|
||
|
||
`merge.wgsl` weights a frame by its tent-filtered share of the label map
|
||
about each pixel (`SeamMap::share`, repeated verbatim), over a window
|
||
`seam_blend_px` wide (64, capped at 4 texels either side). The edge feather
|
||
remains underneath as a factor and, with a 1e-4 floor, as the answer where
|
||
the map names no frame that reaches the pixel. `--feather-only` on
|
||
`examples/merge.rs` merges the old way, for comparison.
|
||
|
||
Known limits: one axis per new frame, so in a multi-row set a frame
|
||
overlapping its left neighbour and the row above is cut along a compromise
|
||
direction; the cost reads grey proxies, so a difference in hue alone is
|
||
invisible to it.
|
||
|
||
## 12. Filling the border instead of cropping it — MI-GAN, read and measured 2026-09-19
|
||
|
||
Raised after the first merges: the ragged border a cylinder leaves could be
|
||
*filled* rather than cropped away. FR-MRG-4 says no boundary fill, on
|
||
§1.3's "not a pixel editor"; this is the evidence for deciding whether to
|
||
revise that, not a revision.
|
||
|
||
**The candidate: MI-GAN** (Sargsyan et al., ICCV 2023, Picsart AI Research).
|
||
Image inpainting designed for mobile: ~6 M parameters, plain convolutions —
|
||
no FFT, no attention — so it quantises to int8 and runs on a phone's DSP,
|
||
with quality close to LaMa and CoModGAN.
|
||
|
||
**Licence: MIT, code and weights alike** (`LICENSE` and `LICENSE-WEIGHTS`
|
||
in the repository, read the same day). The cleanest position of any model
|
||
in the tree — GPL-compatible, store-compatible, no grant to read around.
|
||
|
||
**Export.** The HuggingFace ONNX files are the *pipeline* — uint8 image and
|
||
mask in, crop-around-mask, resize and blend inside the graph, every
|
||
dimension dynamic — and tract refuses them (F6 again). The bare generator
|
||
exports cleanly from the `migan_512_places2.pt` state dict at a fixed
|
||
`1×4×512×512` (`export_migan.py` in the spike directory; the input is
|
||
`mask − 0.5` and the masked RGB in −1..1, the output RGB in −1..1, the
|
||
caller composites). After slimming the graph is **six operator types**:
|
||
`Add, Clip, Conv, LeakyRelu, Mul, Resize`. 28 MB.
|
||
|
||
**Under tract on the reference desktop: loads in 53 ms, runs in 7.4 s per
|
||
512 × 512 tile, f32.** That is the number. The fixture's border is two
|
||
ragged bands across 22 993 px — roughly ninety 512-px tiles at full
|
||
resolution — so a CPU-f32 fill is ten minutes on the desktop and longer on
|
||
the tablet. Three ways to make it viable, none built:
|
||
|
||
1. **Fill at a quarter of the resolution and upsample.** Sky and scree
|
||
tolerate it; twenty-odd tiles, about three minutes on the desktop CPU. A
|
||
background job with the outbox's patience, not an interactive one.
|
||
2. **The tablet's Hexagon through QNN**, where the plain-conv design is the
|
||
point. Measured 2026-10-04 (inference.md §1.5): int8 changes the fill
|
||
(16 dB from f32's), so it ships with 16-bit activations and weights —
|
||
87 ms a tile against 488 ms on the tablet's CPU, the whole graph on the
|
||
NPU, 41 dB from f32 in the hole.
|
||
3. **A WGSL runtime for those six operators.** A project of its own, and
|
||
the only route that would make it interactive on the desktop.
|
||
|
||
Whichever, the fill is a *proposal* under FR-MRG-1's rule — shown, then
|
||
confirmed — and it would sit beside the crop, not replace it: the crop is
|
||
free and honest, the fill is invented pixels, and the photographer chooses.
|
||
|
||
## 13. The fill, built — 2026-09-19, evening
|
||
|
||
Built the same day on `merge/fill`, on the engine (S16) rather than tract,
|
||
and FR-MRG-4 revised to admit it: the border is *cropped or filled*, the
|
||
photographer's choice, the crop the default.
|
||
|
||
**What runs.** `dr_pano::fill` is the engine-independent half: an
|
||
`Inpainter` trait (a 512-px tile in, the same tile out) and `fill_border`,
|
||
which owns everything the model does not — which tiles, what context, how
|
||
to blend. `dr_pano::migan::MiGan` is the trait over the shipped generator
|
||
under `dr_inference_engine` with the new `Role::Inpainter`, so it takes
|
||
whichever rung the device has. The merge job runs the fill at **half the
|
||
composite's resolution**, in a display-ish space (white balance, camera
|
||
matrix, gamma — invertible, so the result goes back to camera-linear and
|
||
into the same linear DNG), and the full-resolution merge samples the fill
|
||
where no frame reached.
|
||
|
||
**What the spike taught, tried in order and kept or dropped.**
|
||
|
||
1. *Context across the coverage edge.* MI-GAN was trained on holes inside
|
||
pictures; given a hole at the picture's edge it invents a structure along
|
||
the open side (white streaks in the sky, on the first try). The known
|
||
content is therefore **mirrored** across the coverage edge into the hole
|
||
and into a 256-px ring, column by column for the top and bottom bands
|
||
and row by row for the sides; the model interpolates between real and
|
||
mirrored sky rather than extrapolating into nothing. *Replicated* rows
|
||
(the edge row continued flat) streaked the grass; a detrended mix (tone
|
||
replicated, texture mirrored) smeared; a low-pass extrapolation banded.
|
||
Mirror stays.
|
||
2. *Coarse to fine.* One pass at the working resolution let the boundary
|
||
leak in — each 512 tile saw only its own corner of the hole. So a
|
||
**coarse pass at a quarter** decides the structure with the whole border
|
||
in a few tiles, and **fine passes in 96-px bands** from the real edge
|
||
outward regenerate texture, each band the only unknown with the previous
|
||
band on its near side and the upsampled coarse fill on its far side.
|
||
3. *The seam.* A hard cut between real and invented showed as a sharpness
|
||
step. The known mask is eroded by a **24-px feather** (48 at half
|
||
resolution) and the fill blended in across that margin by distance to
|
||
the real edge, smoothstep.
|
||
4. *Partial pixels.* The seams were still visible until the cause was found
|
||
upstream of the fill: the camera-space tap stored **black with alpha 1**
|
||
for a pixel the lens correction pushed off the sensor, and the warp
|
||
averaged it in — a dark, poorly interpolated fringe along every frame's
|
||
edge that the fill then continued. `OutputMode::CameraLinear` now
|
||
stores alpha 0 for a pixel that is not there and the merge's warp
|
||
weights by the sampled alpha, so the fringe never enters the composite.
|
||
The mask erosion before the fill dropped from 16 px to 4.
|
||
|
||
5. *What is still wrong, and why it ships anyway.* With the seams gone the
|
||
content itself is the problem in the deep corners: the model, trained
|
||
on Places2, puts bright cloud-and-peak shapes into a sky hole and a
|
||
water-like band under grass — its prior for "top of a picture" and
|
||
"bottom of a landscape", not anything in the context (the same shapes
|
||
appear with the mirror capped, uncapped, and on the CPU as on TensorRT).
|
||
Thin borders are fine; that is most of a hand-held sweep. So the fill
|
||
ships **experimental**: opt-in, previewed, its knobs on the page and
|
||
in the sidecar, and `cargo run -p dr-ui --example fill` re-runs any
|
||
merge's dumped input (`DR_FILL_DUMP=dir`) stage by stage in seconds so
|
||
the next attempt is made from the picture, not from a seven-minute
|
||
merge. Candidates for that attempt: a context that is not a mirror at
|
||
all in deep holes (the coarse pass's own answer, iterated), a sky
|
||
detector that fills sky by extrapolating the gradient and leaves the
|
||
model to texture, or a different model.
|
||
|
||
**Measured, the fixture's twelve frames (22 991 × 5 978), 348 tiles at
|
||
half resolution.** 312 s on ONNX Runtime's CPU pool on the reference
|
||
desktop (≈ 0.8 s a tile). On TensorRT fp16: **100 s**, of which 60 ms a
|
||
tile was the engine hashing the 28 MB model on every acquire (fixed, the
|
||
hash is taken at open) and 150 ms a tile the GPU itself — throttled:
|
||
`trtexec` on the same engine read 23 ms at noon on a cool machine and
|
||
152 ms that evening after two hours of builds, nvidia-smi showing SW power
|
||
cap and thermal slowdown. Cool, the fill is ~10 s. The TensorRT engine
|
||
compiles once, in 13 minutes, cached under the inference directory.
|
||
|
||
**The runtime is a packaging matter.** Arch's `onnxruntime-opt-cuda` has
|
||
no TensorRT provider ("not enabled in this build") and its CUDA provider
|
||
does not load against cuDNN 9, so on this machine the app fell to ORT CPU
|
||
until the official `onnxruntime-linux-x64-gpu_cuda13` tarball (1.30.0,
|
||
which links the system CUDA 13.4 and TensorRT 10.16) was unpacked and
|
||
named with `DARKROOM_ORT_DIR`; `/usr/lib/darkroom` is searched too, for a
|
||
package that ships it. §12's point 2 for the tablet is unchanged.
|
||
|
||
**On the page.** A *Border* choice beside the projection — *Crop to the
|
||
picture* / *Fill the border* — with a caption saying what the fill is; the
|
||
preview re-renders filled when chosen, at preview resolution, so the
|
||
choice is seen before it is confirmed (FR-MRG-1). Greyed out with the reason
|
||
when `migan-512.onnx` is not in the model directory. Under the fill, while
|
||
it is experimental, its six knobs as sliders — working scale, edge
|
||
erosion, coarse pass, band width, mirror depth, seam feather — each
|
||
committing a redraw of the preview. A filled merge's sidecar says `border
|
||
filled` with the knobs used, and its default crop is still the inscribed
|
||
rectangle.
|
||
|
||
**Ships.** `models/inpaint/migan-512.onnx` (LFS, 28 MB, MIT,
|
||
`models/LICENCE.md`), installed by the PKGBUILD and unpacked by the APK
|
||
beside the face and scene models; `tools/export-migan.sh` regenerates it
|
||
from the upstream checkpoint.
|
||
|
||
## 14. The fill, trained — 2026-09-20
|
||
|
||
§13.5 named the remaining fault: in a deep corner the stock model puts its
|
||
Places2 prior — clouds, peaks, a water line — into a hole, because it was
|
||
trained on holes *inside* pictures and a panorama's border is a hole with
|
||
the picture on one side and nothing on the other. Every ring (mirror,
|
||
replicate, detrend) treated the symptom. The fix is a model that has seen
|
||
the real thing: **MI-GAN's 512 generator fine-tuned on border-shaped voids
|
||
cut from the user's own photographs**, in a separate repository
|
||
(`darkroom-infill`, beside this one), so the truth beyond the void is known
|
||
and the model learns one-sided extrapolation.
|
||
|
||
**What it was trained on.** Voids made the way this merge makes them:
|
||
frames with a yaw, a common pitch and per-frame roll, projected onto the
|
||
cylinder and rasterised, the canvas their union's bounding box, the void
|
||
the canvas outside the union — arcs where straight edges bent, cusps where
|
||
frames meet, the bow-tie wedge at a corner (a third of tiles are cut at a
|
||
canvas corner). Voids to 256 px deep at a 512 tile. Half the deep tiles
|
||
train the *second pass*: a no-grad first pass fills the tile, its nearest
|
||
band (64–256 px) is marked known, and the remaining void is the example —
|
||
so the model continues its own output without drift, which is how `fill`
|
||
runs deep voids. Data: ~7 400 pictures — the 1024-px proxy tier of the
|
||
library and ~2 000 raws sampled evenly across every year, developed at
|
||
half size. Loss: hole-weighted L1, VGG16 perceptual, a hinge PatchGAN.
|
||
One night on the reference desktop's RTX 3050.
|
||
|
||
**What changed here.** `FillParams::mirror_depth` **0** is now "no ring":
|
||
the void reaches the tile's edge with nothing beyond, and beyond the band
|
||
being filled the void stays *unknown* rather than presented as known coarse
|
||
fill — the two conditions the model was trained under. Defaults: mirror 0,
|
||
coarse 1 (the coarse pass seeds nothing the model is allowed to see), band
|
||
192. The ring remains on the page for the stock model's sake, at any depth
|
||
above zero. The model file is a drop-in (`models/inpaint/migan-512.onnx`,
|
||
same six operators, same tensors) and the engine loads it unchanged.
|
||
|
||
**Measured, 240 held-out tiles with projection-shaped voids (PSNR in the
|
||
hole, dB / LPIPS on the composite), stock → shipped (step 3 607):** edge
|
||
16.9 → 18.5 / 0.121 → 0.136; corner 14.6 → 16.3 / 0.183 → 0.205; interior
|
||
18.2 → 19.3 / 0.051 → 0.056. Read both columns: the fine-tune gains ~2 dB
|
||
on edges and corners because it stops inventing objects, and *loses* on
|
||
LPIPS because what it paints in a deep void is smoother than the stock
|
||
model's confident wrong texture — LPIPS rewards texture, right or wrong.
|
||
On the fixture's dump at half resolution (the merge's working size) the
|
||
sky corners are sky, with no structure and a faint tone step at worst;
|
||
the ground bands carry a fine texture at the right tone, softer than the
|
||
real scree above them. The stock model's top-left corner on the same
|
||
dump is a glowing invented structure. The training's own record — what
|
||
each loss weighting did, and the two runs abandoned (blur under L1 in
|
||
the hole; a brick pattern under a strong adversarial term against a
|
||
discriminator that had not learned) — is `runs/` in `darkroom-infill`.
|
||
|
||
**Second model, the same day.** The morning's fill was soft in the deep
|
||
ground bands. The afternoon's run trained the generator against MI-GAN's
|
||
own pretrained discriminator (non-saturating loss, lazy R1, feature
|
||
matching), with fresh noise inputs while training and flip/translation
|
||
augmentation of the discriminator's input — both needed, or the generator
|
||
settles into a periodic texture the discriminator cannot see. The shipped
|
||
weights (step 4 750 of `runs/border-v6`) are level with the stock model on
|
||
LPIPS (edge 0.125 / corner 0.187 against 0.121 / 0.183) while keeping the
|
||
PSNR gain (edge 18.0 / corner 15.7 against 16.9 / 14.6). On the fixture
|
||
the ground bands now carry texture at the right tone; at 1:1 a faint
|
||
regular hatch is visible in the deepest part.
|
||
|
||
**What is still wrong.** The hatch, and any deep textured void the
|
||
generator must invent. The better answer for those is not generative:
|
||
seed the void with the picture's own texture in hexagonal cells, let the
|
||
discriminator rank the candidates, and let the generator heal only the
|
||
gaps between cells — built and measured in `darkroom-infill`
|
||
(`infill/hexfill.py`), the most convincing scree corner produced so far,
|
||
and the next thing to port into `dr_pano::fill` (it needs the
|
||
discriminator as a second model, ~80 MB fp16). FR-MRG-4's *experimental*
|
||
stays.
|
||
|
||
## 15. Leaving a frame out, and white clouds — 2026-09-27
|
||
|
||
**A frame is left out from its row, not by starting again.** Until now a
|
||
frame that did not fit ended the job with its name, and the only way on was
|
||
Back, a smaller selection, and every frame read, demosaiced and searched for
|
||
keypoints again. Each row of the Frames table on the merge page now has a
|
||
box, ticked by default. Unticking one leaves the frame out and the rest are
|
||
solved again at once; ticking it brings it back.
|
||
|
||
What makes that cheap is a split in `dr_pano::align`. `match_pairs` does
|
||
the matching and the pairwise RANSAC once over every frame — about 5 s for
|
||
the twelve-frame fixture — and `solve` takes a subset and uses only the
|
||
links among the frames in it, about 0.1 s. Solving a subset by re-aligning
|
||
it from scratch was tried and is wrong: the RANSAC seeds are keyed on frame
|
||
position, so dropping a frame moved every seed after it, and on the fixture
|
||
that was enough to lose a marginal link and strand a neighbour of the frame
|
||
left out. A frame whose only overlap was with one left out is reported
|
||
unaligned, exactly as it would be had it never been measured with it.
|
||
|
||
**A frame that cannot be placed no longer ends the job either** (FR-MRG-5).
|
||
Its row names it and says why, and `Merge` stays off until it is unticked —
|
||
still never a silent drop. The headless example leaves such frames out the
|
||
same way, and takes `--leave-out N` to untick frame `N` once the first
|
||
alignment is in.
|
||
|
||
**Blown highlights stay white.** A clipped photosite reaches the merge as
|
||
camera (1, 1, 1), which the as-shot balance turns magenta. The alignment
|
||
preview balanced it with no highlight rule, so every blown cloud was pink
|
||
on the page. The DNG had a quieter form of the same fault: a frame's gain
|
||
below one moved a blown sample off the white level, and the feather mixed
|
||
it into a neighbour's real sky, after which the develop's own highlight
|
||
desaturation no longer recognised it. The merge shader (`merge.wgsl`) and
|
||
the preview (`grey_if_blown` in `dr_ui::merge`) now write a blown sample,
|
||
before the gain, as the camera value the composite's balance maps to grey —
|
||
the develop pipeline's neutral, fading in from `CLIP_ONSET` exactly as the
|
||
develop's does.
|
||
|
||
**A composite this wide is past two limits, both lifted the same day.**
|
||
They were found on a 22 927 × 8966 panorama from Lightroom, and the fixture's
|
||
own composite (22 993 × 5 980, §11) is past both. rawler's allocation guard,
|
||
sized in samples but worded in pixels, refuses a three-sample DNG past about
|
||
16 700 pixels wide; the copy in `third_party/` carries it raised
|
||
([README](../../third_party/README.md)). And no texture holds such a frame:
|
||
a linear DNG past 8192 pixels now opens on a box-reduced copy, a render finer
|
||
than the copy samples a window of the full resolution, and the export is drawn
|
||
in tiles (FR-DSP-2's note in requirements.md, ARCH §5.3).
|