Files
DarkRoom/docs/panorama.md
T
dtourolle e4b6b6c935 S15.2: XFeat exports at a fixed shape and loads under tract
tools/export-xfeat.sh exports the convolutional network alone at 768×1024
grayscale, on the pattern of export-seg-model.sh: thirteen standard
operator types, no dynamic axes, the keypoint decoding left to Rust.
examples/onnx_probe loads it through the ort-over-tract backend the app
ships with nothing unsupported and runs it in ~300 ms on the desktop CPU.

The weights are Apache-2.0, read from the repository's LICENSE, with no
grant on the checkpoint — recorded in models/LICENCE.md before they land,
as FR-MRG-8 asks. The probe stays: the next model will need the same
check.
2026-09-19 15:24:10 +02:00

14 KiB
Raw Blame History

Panorama

Status: Draft · 2026-09-19 Companion to: requirements.md §3.11 FR-MRG-1 … 11, D18, S15 · architecture.md §5.2, §6.2

The first merge (§3.11): several frames, rotated about one point, become one photograph. This document is how that lands on the pipeline that exists now — which stages, where each runs, how the composite is produced in chunks when it is larger than any texture or any memory, what is ported from where, and what the keypoint model may be under D8.


1. Why it is worth the work

The audience shoots panoramas and leaves the application to stitch them. That is the same workflow break dust was (spot-removal.md §1): a RAW editor that does everything but the one thing, and the photographer's work ends up in a JPEG produced by a tool that never saw the RAW.

It is also the merge whose alignment problem is smallest. A panorama is a rotation — three parameters per frame plus a focal length — with no depth to recover. HDR merge and focus stacking share its data model (D18) and most of its machinery (FR-MRG-3, 5, 6, 7, 10, 11 are written to be general); building the panorama first builds the shared part on the easiest geometry.

2. Non-goals

  • Not structure-from-motion. No translation is solved for. A hand-held set with parallax gets its ghosts hidden by seam placement, and a set with real parallax is not a panorama. COLMAP's front end is the right mental model; its back end is the wrong problem.
  • Not a multi-source Version. D18. The composite is a file, and nothing in the catalog, the sidecar format or sync learns about cross-references.
  • Not boundary fill. Painting pixels that were never captured is the pixel editing §1.3 excludes. Auto-crop is the tool.
  • Not automatic. The tool proposes an alignment and writes nothing until the photographer confirms. Same rule as spot removal and D17, for the same reason: a merge that silently omits or misplaces a frame is the failure this application must not have.
  • Not HDR-panorama in one pass. Until HDR merge exists on its own, a bracketed panorama is bracketed frames merged first, then stitched.

3. What is new, precisely

Nearly all of it, unlike spot removal. The pipeline renders one source to one texture; nothing in the tree detects keypoints, estimates a rotation, warps into a projection, finds a seam, or blends a pyramid. What exists and is reused:

Exists Where Reused for
Render a source to scene-linear on the GPU dr-gpu demosaic → camera profile → working space FR-MRG-2's input, once a tap after lens correction and before tone exists (S15.3)
Tiled rendering with a priority scheduler ARCH §5.3 Pulling source tiles on demand into an output chunk (§5 below)
A non-CFA source entering the pipeline Demosaicer::from_rgba8 The composite's decode path, if the container is a TIFF (S15.1)
DNG matrices read through rawler dr-decode::profile The composite's decode path, if the container is a DNG
Static-shape ONNX under tract, heads decoded in Rust dr-segment The keypoint model (§6)
A batch worker with its own GpuContext, activity row, cancel dr-ui::export FR-MRG-7 verbatim
The 16-bit TIFF encoder with metadata sub-IFDs dr-export::encode FR-MRG-3's writer, extended to linear samples
Multi-select in the grid collections_ui::selected The entry point

New: a core/dr-pano crate holding the geometry (keypoints, matching, the rotation solve), a set of WGSL passes in dr-gpu (reprojection, gain, seam, pyramid blend), the chunked output driver, the container writer, and the dialog.

4. The stages, and where each runs

FR-MRG-10 states the rule; this is the table it was written from.

Stage Cost shape Runs on Why
Source to scene-linear per pixel, full res GPU, the existing pipeline It is the pipeline
Keypoint detection once per frame, at 1024 px CPU, tract (NEON on the tablet) Bounded by frame count, not output size. Same runtime faces and masks use. Hand-written WGSL convolutions for a model that runs five times would be work with no visible gain.
Descriptor matching K² × D per pair CPU, SIMD 2048² × 64 × 10 pairs ≈ 3 GFLOP — tens of milliseconds
Rotation solve, bundle adjustment 3N + 1 parameters, Levenberg–Marquardt CPU Microseconds. Not parallel work.
Preview reprojection per pixel, proxy res GPU, interactive Projection and horizon changes re-warp N proxies at frame rate
Full-resolution warp per output pixel GPU, chunked (§5) The heaviest thing in the application
Gain compensation per overlap region GPU reduction, then N scalars Sums, on the histogram pass's pattern (ARCH §5.5)
Seam finding per overlap pixel GPU-friendly variant Graph cut resists the GPU; a distance-transform or per-column DP seam does not. The algorithm is chosen for the GPU, not for the paper.
Multi-band blend per pixel × levels GPU, chunked Laplacian pyramids are separable convolutions — the detail stage's shape
Encode per pixel, once CPU, streamed per chunk row As export does

5. Chunked in output space

FR-MRG-11 forbids holding the composite as one texture, and two facts force it before memory does:

  • max_texture_dimension_2d is 8192 on many mobile GPUs and 16384 on desktop. A three-row panorama is routinely 20 000 px wide.
  • Five 24 MP frames at working precision are ~1 GB together. The tablet does not have it.

The geometry is known before any full-resolution pixel exists. Alignment runs on proxies; what comes out is a rotation per frame, a focal length, a projection and an output rectangle. From those, every output pixel's source coordinates in every frame are a closed-form function. That is what makes chunking simple rather than clever:

for each output chunk C (e.g. 2048 × 2048, in output space):
    frames_in(C) = frames whose projected footprint intersects C
    for each frame F in frames_in(C):
        source tiles T(F, C) = tiles of F that project into C, plus a margin
        render T(F, C) to scene-linear through the pipeline's tile cache
        warp T(F, C) into C's coordinate frame           ← GPU
    gain-correct, seam, blend within C                     ← GPU, with overlap
    read C back, encode its rows                           ← CPU, streamed

The working set is one chunk, its per-frame warped copies, and the source tiles that fed them. It does not grow with the composite.

The blend needs a margin. A Laplacian pyramid of L levels reads 2^L pixels beyond the chunk edge; a chunk is therefore rendered with a margin of that width and the margin discarded after the blend. Seams cross chunk boundaries and must agree on both sides: the seam is found once at a reduced resolution over the whole overlap (which fits — it is a mask, not an image), then upsampled into each chunk. The same is true of gain: the scalars are solved once from proxy-resolution overlaps and applied everywhere.

Source tiles are the pipeline's tiles. ARCH §5.3's cache keys by (VersionId, tile, zoom, graph_hash_prefix); the merge asks for tiles of a neutral graph at zoom 1 and gets the same caching every other consumer does. A tile pulled for one chunk is usually needed by the neighbouring chunk, and stays hot for it.

6. The keypoint model

FR-MRG-8: works without weights, better with them. The licence read comes first (D13's lesson, S15.2).

Model Licence Fits tract? Position
XFeat (CVPR 2024) Apache-2.0 Plain convolutions, fully convolutional, the repo ships an ONNX export Chosen. Fixed 1024 px input, dense heatmap and descriptor map out, NMS and top-K in Rust — the yolo26 pattern
DISK Apache-2.0 U-Net, static Second choice; stronger descriptors, ~3–4× the compute
ALIKE BSD-3 Plain convolutions Fallback if XFeat's export fails F6
ALIKED BSD-3 Deformable convolution in the descriptor head Unlikely to load
SuperPoint, SuperGlue, R2D2, SiLK, MASt3R non-commercial — Out on licence
LightGlue Apache-2.0 Transformer over a variable keypoint count Not until mutual-nearest-neighbour matching fails on a real set

S15.2, 2026-09-19: XFeat loads under tract. tools/export-xfeat.sh exports the network alone at 768×1024 — thirteen operator types, all standard: Conv, InstanceNormalization, AveragePool, Resize, Slice, Transpose, Reshape, Concat, Add, Relu, Sigmoid, ReduceMean, Unsqueeze — and examples/onnx_probe.rs loads the 2.8 MB file through the app's own ort-over-tract backend with nothing unsupported, in 28 ms, and runs it in ~300 ms on the reference desktop's CPU. The weights ship as models/keypoints/xfeat-1024.onnx, recorded in models/LICENCE.md. Still to do: the tablet figure (S15.4), and a keypoint-level comparison against the PyTorch reference once the Rust decoder exists — the probe proves the graph runs, not that the numbers match.

The outputs are three maps at 1/8 resolution, 96×128 for the export size: 64-channel descriptors, 65-channel keypoint logits (each 8×8 cell's position plus "none"), and a reliability heatmap. The Rust decoder is: softmax over the 65, pixel-shuffle the first 64 to full resolution, 5×5 non-maximum suppression, top-k by reliability, bilinear sampling of the descriptor at each keypoint, L2 normalise. That is detectAndCompute in the reference, minus the network.

Without weights: AKAZE (BSD, akaze from rust-cv), which is adequate on well-textured overlaps and worse on sky, repeated structure and exposure drift — which is where a learned detector earns its place.

Matching is mutual nearest neighbour with a ratio test, then RANSAC on a rotation model. For a panorama — one lens, near-pure rotation, 20–40 % overlap — that is what Hugin and OpenCV's stitcher use, and it is enough.

7. What is ported from where

Nothing is linked; everything is read.

Source Licence Taken
OpenCV modules/stitching Apache-2.0 The stage layout — Brown & Lowe (2007) as a set of small classes with one job each — and the warpers' projection maths
OpenPano (ppwwyyxx) MIT (verify on read) The estimation and bundle-adjustment maths, function by function, with outputs diffed against it
enblend-enfuse GPLv2+ Seam-line optimisation and Burt–Adelson multi-band blending
Hugin nona GPLv2+ The GLSL remapper, as the reference for the WGSL warp

The golden set (§8 of the requirements) is OpenCV's stitcher on the same inputs: a reference output to compare against, within a tolerance calibrated the way S9 calibrates R1.

8. The output file

FR-MRG-3. Scene-linear, ≥ 16 bits, wide gamut, the first source's capture metadata, named from the first source with a -pano suffix, beside it.

Two containers were candidates and S15.1 decided, on 2026-09-19:

  • Linear DNG. PhotometricInterpretation = LinearRaw, three samples per pixel, ColorMatrix1 carried from the first source. Re-enters through rawler as Format::Dng with no new decode path, if rawler reads it back. What Lightroom writes.
  • Float TIFF. SampleFormat = IEEEFP, 16 or 32 bits, an ICC profile for the working space. Needs Format::Tiff and a decode path, but the writer is the existing encoder with a different sample type, and nothing about it is uncertain.

Linear DNG. examples/linear_dng.rs hand-rolls a 64 × 48 LinearRaw DNG — one IFD, uncompressed 16-bit RGB, DNGVersion, ColorMatrix1, AsShotNeutral, CalibrationIlluminant1 — and rawler 0.7 reads it back: cpp 3, the samples interleaved as written, the matrix parsed into the camera definition, and CameraProfile::extract builds the same profile it would for a camera file. ImageMagick's libraw reads the same bytes. What does not yet work is dr_decode::decode, which accepts the file as CFA and hands the pipeline three times the samples it expects: the cpp == 3 branch is the work, and it is the only decode work.

The composite therefore enters the pipeline as a non-CFA, linear source — from_rgba8's sibling with non_linear = false and the colour matrix carried from the DNG — and is developed as any RAW is. The writer is the example's IFD, grown up: tiled rather than one strip (FR-MRG-11 encodes per chunk), and carrying the first source's EXIF in a sub-IFD as dr-export already does.

9. Interaction

  • The entry is the grid's selection: two or more images, one action, "Merge to panorama". One image, or images from different roots, and the action says why it is unavailable.
  • The dialog shows the aligned proxies in the chosen projection, with the projection, horizon and crop controls of FR-MRG-4, and the per-frame residuals. A frame that failed to align is named there (FR-MRG-5), and the merge cannot be confirmed with it in the set.
  • Confirm starts the FR-MRG-7 job. The composite appears in the grid when the file is written and catalogued, beside its sources, with the merge as the first entry in its history.

10. Order of work

  1. S15, all four, before anything else. (1) and (2) are a day each and either can change the design.
  2. dr-pano: keypoints (AKAZE first, XFeat when S15.2 passes), matching, RANSAC, rotation solve. Unit-tested against synthetic rotations of one frame, where the answer is known exactly.
  3. The working-space tap, and the preview reprojection pass. At this point the dialog can show an alignment.
  4. The chunked driver with a feathered blend — the whole path end to end, writing a file, before the blend is good.
  5. Gain, seams, multi-band.
  6. The container, the catalog entry, provenance, the history entry.
  7. Tablet: NFR-MRG-1's figure, and FR-MRG-9's ceiling.