Files
DarkRoom/docs/panorama.md
T
dtourolle 67f225beba panorama.md: MI-GAN as the border filler — MIT, six operators, 7.4 s a tile
Read and measured, not built. The bare 512 generator exports at a fixed
shape and loads under tract with nothing unsupported; at f32 on the
desktop CPU it takes 7.4 s per 512×512 tile, which puts a full-resolution
fill of the fixture's border at ten minutes. The three routes that would
make it viable are recorded, with the quarter-resolution fill the cheapest
and Hexagon int8 the one the model was designed for.
2026-09-19 15:30:49 +02:00

391 lines
22 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Panorama
**Status:** Draft · 2026-09-19
**Companion to:** [requirements.md](requirements.md) §3.11 FR-MRG-1 … 11, D18, S15 · [architecture.md](architecture.md) §5.2, §6.2
The first merge (§3.11): several frames, rotated about one point, become one
photograph. This document is how that lands on the pipeline that exists now —
which stages, where each runs, how the composite is produced in chunks when it
is larger than any texture or any memory, what is ported from where, and what
the keypoint model may be under D8.
---
## 1. Why it is worth the work
The audience shoots panoramas and leaves the application to stitch them. That
is the same workflow break dust was (`spot-removal.md` §1): a RAW editor that
does everything but the one thing, and the photographer's work ends up in a
JPEG produced by a tool that never saw the RAW.
It is also the merge whose alignment problem is smallest. A panorama is a
rotation — three parameters per frame plus a focal length — with no depth to
recover. HDR merge and focus stacking share its data model (D18) and most of
its machinery (FR-MRG-3, 5, 6, 7, 10, 11 are written to be general); building
the panorama first builds the shared part on the easiest geometry.
## 2. Non-goals
- **Not structure-from-motion.** No translation is solved for. A hand-held set
with parallax gets its ghosts hidden by seam placement, and a set with real
parallax is not a panorama. COLMAP's front end is the right mental model;
its back end is the wrong problem.
- **Not a multi-source Version.** D18. The composite is a file, and nothing in
the catalog, the sidecar format or sync learns about cross-references.
- **Not boundary fill.** Painting pixels that were never captured is the pixel
editing §1.3 excludes. Auto-crop is the tool.
- **Not automatic.** The tool proposes an alignment and writes nothing until
the photographer confirms. Same rule as spot removal and D17, for the same
reason: a merge that silently omits or misplaces a frame is the failure this
application must not have.
- **Not HDR-panorama in one pass.** Until HDR merge exists on its own, a
bracketed panorama is bracketed frames merged first, then stitched.
## 3. What is new, precisely
Nearly all of it, unlike spot removal. The pipeline renders one source to one
texture; nothing in the tree detects keypoints, estimates a rotation, warps
into a projection, finds a seam, or blends a pyramid. What exists and is
reused:
| Exists | Where | Reused for |
|---|---|---|
| Render a source through the fused pass, with a linear f16 output mode | `dr-gpu` demosaic → `AdjustPass`, `OutputMode::LinearWorking` | FR-MRG-2's camera-space input, as a compose entry with no operations and the profile uniforms neutral (S15.3) |
| Tiled rendering with a priority scheduler | ARCH §5.3 | Pulling source tiles on demand into an output chunk (§5 below) |
| A non-CFA source entering the pipeline | `Demosaicer::from_rgba8` | The composite's decode path, if the container is a TIFF (S15.1) |
| DNG matrices read through rawler | `dr-decode::profile` | The composite's decode path, if the container is a DNG |
| Static-shape ONNX under tract, heads decoded in Rust | `dr-segment` | The keypoint model (§6) |
| A batch worker with its own `GpuContext`, activity row, cancel | `dr-ui::export` | FR-MRG-7 verbatim |
| The 16-bit TIFF encoder with metadata sub-IFDs | `dr-export::encode` | FR-MRG-3's writer, extended to linear samples |
| Multi-select in the grid | `collections_ui::selected` | The entry point |
New: a `core/dr-pano` crate holding the geometry (keypoints, matching, the
rotation solve), a set of WGSL passes in `dr-gpu` (reprojection, gain,
seam, pyramid blend), the chunked output driver, the container writer, and
the dialog.
## 4. The stages, and where each runs
FR-MRG-10 states the rule; this is the table it was written from.
| Stage | Cost shape | Runs on | Why |
|---|---|---|---|
| Source to camera-linear | per pixel, full res | GPU, the existing pipeline | It *is* the pipeline, stopped early |
| Keypoint detection | once per frame, at 1024 px | CPU, tract (NEON on the tablet) | Bounded by frame count, not output size. Same runtime faces and masks use. Hand-written WGSL convolutions for a model that runs five times would be work with no visible gain. |
| Descriptor matching | K² × D per pair | CPU, SIMD | 2048² × 64 × 10 pairs ≈ 3 GFLOP — tens of milliseconds |
| Rotation solve, bundle adjustment | 3N + 1 parameters, Levenberg–Marquardt | CPU | Microseconds. Not parallel work. |
| Preview reprojection | per pixel, proxy res | GPU, interactive | Projection and horizon changes re-warp N proxies at frame rate |
| Full-resolution warp | per output pixel | GPU, chunked (§5) | The heaviest thing in the application |
| Gain compensation | per overlap region | GPU reduction, then N scalars | Sums, on the histogram pass's pattern (ARCH §5.5) |
| Seam finding | per overlap pixel | GPU-friendly variant | Graph cut resists the GPU; a distance-transform or per-column DP seam does not. The algorithm is chosen for the GPU, not for the paper. |
| Multi-band blend | per pixel × levels | GPU, chunked | Laplacian pyramids are separable convolutions — the detail stage's shape |
| Encode | per pixel, once | CPU, streamed per chunk row | As export does |
## 5. Chunked in output space
FR-MRG-11 forbids holding the composite as one texture, and two facts force it
before memory does:
- `max_texture_dimension_2d` is 8192 on many mobile GPUs and 16384 on desktop.
A three-row panorama is routinely 20 000 px wide.
- Five 24 MP frames at working precision are ~1 GB together. The tablet does
not have it.
**The geometry is known before any full-resolution pixel exists.** Alignment
runs on proxies; what comes out is a rotation per frame, a focal length, a
projection and an output rectangle. From those, every output pixel's source
coordinates in every frame are a closed-form function. That is what makes
chunking simple rather than clever:
```
for each output chunk C (e.g. 2048 × 2048, in output space):
frames_in(C) = frames whose projected footprint intersects C
for each frame F in frames_in(C):
source tiles T(F, C) = tiles of F that project into C, plus a margin
render T(F, C) to scene-linear through the pipeline's tile cache
warp T(F, C) into C's coordinate frame ← GPU
gain-correct, seam, blend within C ← GPU, with overlap
read C back, encode its rows ← CPU, streamed
```
The working set is one chunk, its per-frame warped copies, and the source
tiles that fed them. It does not grow with the composite.
**The blend needs a margin.** A Laplacian pyramid of L levels reads
2^L pixels beyond the chunk edge; a chunk is therefore rendered with a margin
of that width and the margin discarded after the blend. Seams cross chunk
boundaries and must agree on both sides: the seam is found once at a reduced
resolution over the whole overlap (which fits — it is a mask, not an image),
then upsampled into each chunk. The same is true of gain: the scalars are
solved once from proxy-resolution overlaps and applied everywhere.
**Source tiles are the pipeline's tiles.** ARCH §5.3's cache keys by
`(VersionId, tile, zoom, graph_hash_prefix)`; the merge asks for tiles of a
neutral graph at zoom 1 and gets the same caching every other consumer does.
A tile pulled for one chunk is usually needed by the neighbouring chunk, and
stays hot for it.
### 5.1 The tap — S15.3, answered by reading the composer
The fused shader's order, fixed by `operation.rs`'s own tests: warp → as-shot
white balance → operations → base curve → camera matrix → store. The store is
either the display encode or, in `OutputMode::LinearWorking`, an unclipped
`rgba16float` of linear sRGB. That mode exists for the detail stage and is
selected from the operations, never by a caller flag, so that a shader and
the texture bound to it cannot disagree.
The merge wants the values *before* the curve and matrix (FR-MRG-2), and the
composer already makes that a matter of uniforms rather than structure: the
white balance, the matrix and the curve's active flag are all in the reserved
uniform block, and a fused pass with no operations, `as_shot_wb = 1`,
`cam_to_srgb = I` and `base_curve_last.z = 0` stores exactly camera-linear
RGB after the warp. So the tap is:
- `EditGraph::compose_camera_linear()` — the `LinearWorking` tail with an
empty operation list and identity framing, paired by name with
- `AdjustPass::render_camera_linear()` — binds the f16 target, fills the
reserved uniforms neutral instead of from the source, returns the texture,
- and a float readback beside the existing 8-bit one.
Nothing in the chain moves. **Precision:** the tap and every chunk buffer
after it should be `rgba32float`, not f16. A 14-bit sensor has 16 384 steps
to white; f16 has 2 048 in the top octave, and a composite that is going to
be re-developed deserves the sensor's precision. The cost is 2× on buffers
FR-MRG-11 already bounds.
**What the DNG carries as a consequence:** the first source's `Make`,
`Model` and `UniqueCameraModel` — so `base_curve::for_body` finds the 6D's
curve — its `ColorMatrix1`/`2` with illuminants, and its `AsShotNeutral`. The
composite then develops through the same profile as its sources, applied
once. The spike's 64 × 48 file (§8) already carries the matrix and neutral;
the body name is a string.
## 6. The keypoint model
FR-MRG-8: works without weights, better with them. The licence read comes
first (D13's lesson, S15.2).
| Model | Licence | Fits tract? | Position |
|---|---|---|---|
| **XFeat** (CVPR 2024) | Apache-2.0 | Plain convolutions, fully convolutional, the repo ships an ONNX export | **Chosen.** Fixed 1024 px input, dense heatmap and descriptor map out, NMS and top-K in Rust — the yolo26 pattern |
| DISK | Apache-2.0 | U-Net, static | Second choice; stronger descriptors, ~3–4× the compute |
| ALIKE | BSD-3 | Plain convolutions | Fallback if XFeat's export fails F6 |
| ALIKED | BSD-3 | Deformable convolution in the descriptor head | Unlikely to load |
| SuperPoint, SuperGlue, R2D2, SiLK, MASt3R | non-commercial | — | Out on licence |
| LightGlue | Apache-2.0 | Transformer over a variable keypoint count | Not until mutual-nearest-neighbour matching fails on a real set |
**S15.2, 2026-09-19: XFeat loads under tract.** `tools/export-xfeat.sh`
exports the network alone at 768×1024 — thirteen operator types, all
standard: `Conv`, `InstanceNormalization`, `AveragePool`, `Resize`, `Slice`,
`Transpose`, `Reshape`, `Concat`, `Add`, `Relu`, `Sigmoid`, `ReduceMean`,
`Unsqueeze` — and
[`examples/onnx_probe.rs`](../core/dr-segment/examples/onnx_probe.rs) loads
the 2.8 MB file through the app's own `ort`-over-tract backend with nothing
unsupported, in 28 ms, and runs it in **~300 ms on the reference desktop's
CPU**. The weights ship as `models/keypoints/xfeat-1024.onnx`, recorded in
`models/LICENCE.md`. **On the tablet** (S15.4's CPU half, same day):
`tools/onnx-probe-on-device.sh` cross-builds the probe, and the same file
runs in **~400 ms per frame** on the reference tablet's NEON cores (ROD2-W09,
SM8635), with output ranges identical to the desktop's — inside NFR-MRG-1's
1 s per frame with room to spare, and 1.3× the desktop rather than the 2×
faces.md §9 measured for its scan. Still to do: a keypoint-level comparison
against the PyTorch reference once the Rust decoder exists — the probe
proves the graph runs, not that the numbers match.
The outputs are three maps at 1/8 resolution, 96×128 for the export size:
64-channel descriptors, 65-channel keypoint logits (each 8×8 cell's position
plus "none"), and a reliability heatmap. The Rust decoder is: softmax over the
65, pixel-shuffle the first 64 to full resolution, 5×5 non-maximum
suppression, top-k by reliability, bilinear sampling of the descriptor at
each keypoint, L2 normalise. That is `detectAndCompute` in the reference,
minus the network.
Without weights: AKAZE (BSD, `akaze` from rust-cv), which is adequate on
well-textured overlaps and worse on sky, repeated structure and exposure
drift — which is where a learned detector earns its place.
Matching is mutual nearest neighbour with a ratio test, then RANSAC on a
rotation model. For a panorama — one lens, near-pure rotation, 20–40 %
overlap — that is what Hugin and OpenCV's stitcher use, and it is enough.
## 7. What is ported from where
Nothing is linked; everything is read.
| Source | Licence | Taken |
|---|---|---|
| OpenCV `modules/stitching` | Apache-2.0 | The stage layout — Brown & Lowe (2007) as a set of small classes with one job each — and the warpers' projection maths |
| OpenPano (ppwwyyxx) | MIT (verify on read) | The estimation and bundle-adjustment maths, function by function, with outputs diffed against it |
| enblend-enfuse | GPLv2+ | Seam-line optimisation and Burt–Adelson multi-band blending |
| Hugin `nona` | GPLv2+ | The GLSL remapper, as the reference for the WGSL warp |
The golden set (§8 of the requirements) is OpenCV's stitcher on the same
inputs: a reference output to compare against, within a tolerance calibrated
the way S9 calibrates R1.
## 8. The output file
FR-MRG-3. A linear DNG at the source's native scale: `u16` samples on the
first source's black-subtracted scale, `WhiteLevel` = its white minus its
black (13 023 for the 6D set: 15 070 − 2 047), `BlackLevel` = 0. Not rescaled
to 65 535 — the sensor had 14 bits and the file says so, and a value the
sensor could not have produced is not invented by a multiply. The first
source's `Make`, `Model`, `UniqueCameraModel`, `ColorMatrix1/2`,
`CalibrationIlluminant1/2`, `AsShotNeutral` and EXIF are carried, so the
composite develops through the same profile as its sources. Named from the
first source with a `-pano` suffix, beside it.
Three samples per pixel rather than a CFA: the warp resamples, and there is no
sensor grid to mosaic back onto. Nothing else about being a RAW is lost —
no white balance, no curve, no matrix, no clip has been applied — and the
photographer develops the panorama afterwards as one photograph.
The sources are portrait frames in the 6D set: `Orientation` is applied
before alignment (learned features are not rotation-invariant) and the
composite is written upright with `Orientation = 1`.
Two containers were candidates and S15.1 decided, on 2026-09-19:
- **Linear DNG.** `PhotometricInterpretation = LinearRaw`, three samples per
pixel, `ColorMatrix1` carried from the first source. Re-enters through
rawler as `Format::Dng` with no new decode path, *if* rawler reads it back.
What Lightroom writes.
- **Float TIFF.** `SampleFormat = IEEEFP`, 16 or 32 bits, an ICC profile for
the working space. Needs `Format::Tiff` and a decode path, but the writer is
the existing encoder with a different sample type, and nothing about it is
uncertain.
**Linear DNG.** [`examples/linear_dng.rs`](../core/dr-decode/examples/linear_dng.rs)
hand-rolls a 64 × 48 `LinearRaw` DNG — one IFD, uncompressed 16-bit RGB,
`DNGVersion`, `ColorMatrix1`, `AsShotNeutral`, `CalibrationIlluminant1` — and
rawler 0.7 reads it back: `cpp 3`, the samples interleaved as written, the
matrix parsed into the camera definition, and `CameraProfile::extract` builds
the same profile it would for a camera file. ImageMagick's libraw reads the
same bytes. What does *not* yet work is `dr_decode::decode`, which accepts the
file as CFA and hands the pipeline three times the samples it expects: the
`cpp == 3` branch is the work, and it is the only decode work.
The composite therefore enters the pipeline as a non-CFA, *linear* source —
`from_rgba8`'s sibling with `non_linear = false` and the colour matrix carried
from the DNG — and is developed as any RAW is. The writer is the example's
IFD, grown up: tiled rather than one strip (FR-MRG-11 encodes per chunk), and
carrying the first source's EXIF in a sub-IFD as `dr-export` already does.
## 9. Interaction
- The entry is the grid's selection: two or more images, one action, "Merge
to panorama". One image, or images from different roots, and the action
says why it is unavailable.
- The dialog shows the aligned proxies in the chosen projection, with the
projection, horizon and crop controls of FR-MRG-4, and the per-frame
residuals. A frame that failed to align is named there (FR-MRG-5), and the
merge cannot be confirmed with it in the set.
- Confirm starts the FR-MRG-7 job. The composite appears in the grid when the
file is written and catalogued, beside its sources, with the merge as the
first entry in its history.
## 10. Order of work
1. **S15**, all four, before anything else. (1) and (2) are a day each and
either can change the design.
2. `dr-pano`: keypoints (AKAZE first, XFeat when S15.2 passes), matching,
RANSAC, rotation solve. Unit-tested against synthetic rotations of one
frame, where the answer is known exactly.
3. The working-space tap, and the preview reprojection pass. At this point the
dialog can show an alignment.
4. The chunked driver with a feathered blend — the whole path end to end,
writing a file, before the blend is good.
5. Gain, seams, multi-band.
6. The container, the catalog entry, provenance, the history entry.
7. Tablet: NFR-MRG-1's figure, and FR-MRG-9's ceiling.
## 11. Where it stands — 2026-09-19, end of the first day
Built, on branch `merge/panorama`, in the order §10 gave:
| Piece | Where | State |
|---|---|---|
| Geometry: keypoints, matching, homography, focal, bundle adjustment, projections | `core/dr-pano` | Done; 33 tests without a model; the fixture aligns in 4.5 s |
| XFeat at two shapes under tract | `models/keypoints`, `dr_pano::xfeat` | Done; 300 ms/frame desktop, 400 ms tablet |
| The camera-space tap | `OutputMode::CameraLinear`, `AdjustPass::render_camera_linear` | Done, `rgba32float`, tiles by view rect |
| Linear DNG writer, streamed | `dr_export::write_linear_dng` | Done; rawler reads it back |
| A three-sample `RawImage` re-entering the pipeline | `dr-decode`, `DemosaicedImage::from_linear_rgb16` | Done |
| Warp, accumulate, resolve, chunk by chunk | `dr_gpu::MergePass`, `merge.wgsl` | Done; feathered blend, scalar gain |
| The job: load, proxies, align, gains, confirm, merge, provenance | `dr_ui::merge` | Done; `examples/merge.rs` drives it headless |
| The page: table, preview, projection, Merge/Stop/Back; the grid's button | `merge.slint`, `merge_ui.rs` | Done; `DARKROOM_START_MERGE=a.CR2,b.CR2` lands on it |
| Placement beside the sources through the outbox, rescan | `merge_ui.rs` | Done, untested against a server |
**Measured on the fixture (desktop, 12 × 20 MP, Intel adapter):** proxies
and keypoints 4 s, alignment 4.5–12.6 s (load-sensitive: the matcher is
every core), the merge **26 s for a 22 993 × 5 980 composite** in twelve
bands of 2048 × 512 chunks, 45 s all told, an 825 MB DNG. NFR-MRG-1's 60 s
holds on the desktop with room; the tablet's figure is still S15.4's open
half.
**Open, in the order they matter:**
1. **Auto-crop (FR-MRG-4).** The merge returns a coverage mask per band and
the file carries the black border. The largest inscribed rectangle over
the coverage, then the DNG's `DefaultCropOrigin`/`DefaultCropSize`, so
nothing is thrown away and the develop view opens on the picture.
2. **Seams and the pyramid** (§10 step 5). The feather hides exposure and
small misalignment; parallax on the near slope will show as a soft
double edge at 1:1.
3. **Vignetting in the tap.** The lens profile's distortion is applied
before the fetch; its vignetting is an operation and is not. Frame edges
are darker than their centres by the lens's falloff, and the feather
averages them into the overlaps.
4. **The tablet:** memory (twelve 40 MB sensor buffers on the CPU, one
demosaiced frame at a time on the GPU), the figure, and FR-MRG-9's
ceiling.
5. **Horizon and drag-to-correct (FR-MRG-4, the proposed 4a).** The
alignment failed on nothing in the fixture; the interaction waits for a
set it fails on.
6. **`derived_from` names sources by file name**, not content hash: the
catalog's `content_hash` is null for most images most of the time. The
hash can join it when the catalog has one.
## 12. Filling the border instead of cropping it — MI-GAN, read and measured 2026-09-19
Raised after the first merges: the ragged border a cylinder leaves could be
*filled* rather than cropped away. FR-MRG-4 says no boundary fill, on
§1.3's "not a pixel editor"; this is the evidence for deciding whether to
revise that, not a revision.
**The candidate: MI-GAN** (Sargsyan et al., ICCV 2023, Picsart AI Research).
Image inpainting designed for mobile: ~6 M parameters, plain convolutions —
no FFT, no attention — so it quantises to int8 and runs on a phone's DSP,
with quality close to LaMa and CoModGAN.
**Licence: MIT, code and weights alike** (`LICENSE` and `LICENSE-WEIGHTS`
in the repository, read the same day). The cleanest position of any model
in the tree — GPL-compatible, store-compatible, no grant to read around.
**Export.** The HuggingFace ONNX files are the *pipeline* — uint8 image and
mask in, crop-around-mask, resize and blend inside the graph, every
dimension dynamic — and tract refuses them (F6 again). The bare generator
exports cleanly from the `migan_512_places2.pt` state dict at a fixed
`1×4×512×512` (`export_migan.py` in the spike directory; the input is
`mask − 0.5` and the masked RGB in −1..1, the output RGB in −1..1, the
caller composites). After slimming the graph is **six operator types**:
`Add, Clip, Conv, LeakyRelu, Mul, Resize`. 28 MB.
**Under tract on the reference desktop: loads in 53 ms, runs in 7.4 s per
512 × 512 tile, f32.** That is the number. The fixture's border is two
ragged bands across 22 993 px — roughly ninety 512-px tiles at full
resolution — so a CPU-f32 fill is ten minutes on the desktop and longer on
the tablet. Three ways to make it viable, none built:
1. **Fill at a quarter of the resolution and upsample.** Sky and scree
tolerate it; twenty-odd tiles, about three minutes on the desktop CPU. A
background job with the outbox's patience, not an interactive one.
2. **int8 on the tablet's Hexagon through QNN**, where the plain-conv design
is the point and the whole graph should run in milliseconds. The setup
exists from the eye-state work; MI-GAN is a candidate for the same path.
3. **A WGSL runtime for those six operators.** A project of its own, and
the only route that would make it interactive on the desktop.
Whichever, the fill is a *proposal* under FR-MRG-1's rule — shown, then
confirmed — and it would sit beside the crop, not replace it: the crop is
free and honest, the fill is invented pixels, and the photographer chooses.