docs/ had 26 developer documents flat beside the manual, and the two audiences are very differently sized: most readers want the manual and the gesture reference, a few want the register, the designs and the measurements. The manual and gestures.md stay at the top; everything for someone changing the code moves to docs/dev/, and the two documents that name their own successors — the v0.1 milestone and the UI-refinement plan — go to docs/dev/archive/ rather than being deleted, since both are still cited. docs/README.md is the index, users first. Every reference follows: code comments, Cargo manifests, the workflows, the pre-commit hook, the bench and traceability tools (which locate the repo root by docs/dev/requirements.md now), packaging, the Docker READMEs, CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level deeper and is regenerated. Links out of the moved documents into the tree gain a level; a link checker over every Markdown file finds none broken.
547 lines
32 KiB
Markdown
547 lines
32 KiB
Markdown
# Panorama
|
||
|
||
**Status:** Draft · 2026-09-19
|
||
**Companion to:** [requirements.md](requirements.md) §3.11 FR-MRG-1 … 11, D18, S15 · [architecture.md](architecture.md) §5.2, §6.2
|
||
|
||
The first merge (§3.11): several frames, rotated about one point, become one
|
||
photograph. This document is how that lands on the pipeline that exists now —
|
||
which stages, where each runs, how the composite is produced in chunks when it
|
||
is larger than any texture or any memory, what is ported from where, and what
|
||
the keypoint model may be under D8.
|
||
|
||
---
|
||
|
||
## 1. Why it is worth the work
|
||
|
||
The audience shoots panoramas and leaves the application to stitch them. That
|
||
is the same workflow break dust was (`spot-removal.md` §1): a RAW editor that
|
||
does everything but the one thing, and the photographer's work ends up in a
|
||
JPEG produced by a tool that never saw the RAW.
|
||
|
||
It is also the merge whose alignment problem is smallest. A panorama is a
|
||
rotation — three parameters per frame plus a focal length — with no depth to
|
||
recover. HDR merge and focus stacking share its data model (D18) and most of
|
||
its machinery (FR-MRG-3, 5, 6, 7, 10, 11 are written to be general); building
|
||
the panorama first builds the shared part on the easiest geometry.
|
||
|
||
## 2. Non-goals
|
||
|
||
- **Not structure-from-motion.** No translation is solved for. A hand-held set
|
||
with parallax gets its ghosts hidden by seam placement, and a set with real
|
||
parallax is not a panorama. COLMAP's front end is the right mental model;
|
||
its back end is the wrong problem.
|
||
- **Not a multi-source Version.** D18. The composite is a file, and nothing in
|
||
the catalog, the sidecar format or sync learns about cross-references.
|
||
- **Not boundary fill.** Painting pixels that were never captured is the pixel
|
||
editing §1.3 excludes. Auto-crop is the tool.
|
||
- **Not automatic.** The tool proposes an alignment and writes nothing until
|
||
the photographer confirms. Same rule as spot removal and D17, for the same
|
||
reason: a merge that silently omits or misplaces a frame is the failure this
|
||
application must not have.
|
||
- **Not HDR-panorama in one pass.** Until HDR merge exists on its own, a
|
||
bracketed panorama is bracketed frames merged first, then stitched.
|
||
|
||
## 3. What is new, precisely
|
||
|
||
Nearly all of it, unlike spot removal. The pipeline renders one source to one
|
||
texture; nothing in the tree detects keypoints, estimates a rotation, warps
|
||
into a projection, finds a seam, or blends a pyramid. What exists and is
|
||
reused:
|
||
|
||
| Exists | Where | Reused for |
|
||
|---|---|---|
|
||
| Render a source through the fused pass, with a linear f16 output mode | `dr-gpu` demosaic → `AdjustPass`, `OutputMode::LinearWorking` | FR-MRG-2's camera-space input, as a compose entry with no operations and the profile uniforms neutral (S15.3) |
|
||
| Tiled rendering with a priority scheduler | ARCH §5.3 | Pulling source tiles on demand into an output chunk (§5 below) |
|
||
| A non-CFA source entering the pipeline | `Demosaicer::from_rgba8` | The composite's decode path, if the container is a TIFF (S15.1) |
|
||
| DNG matrices read through rawler | `dr-decode::profile` | The composite's decode path, if the container is a DNG |
|
||
| Static-shape ONNX under tract, heads decoded in Rust | `dr-segment` | The keypoint model (§6) |
|
||
| A batch worker with its own `GpuContext`, activity row, cancel | `dr-ui::export` | FR-MRG-7 verbatim |
|
||
| The 16-bit TIFF encoder with metadata sub-IFDs | `dr-export::encode` | FR-MRG-3's writer, extended to linear samples |
|
||
| Multi-select in the grid | `collections_ui::selected` | The entry point |
|
||
|
||
New: a `core/dr-pano` crate holding the geometry (keypoints, matching, the
|
||
rotation solve), a set of WGSL passes in `dr-gpu` (reprojection, gain,
|
||
seam, pyramid blend), the chunked output driver, the container writer, and
|
||
the dialog.
|
||
|
||
## 4. The stages, and where each runs
|
||
|
||
FR-MRG-10 states the rule; this is the table it was written from.
|
||
|
||
| Stage | Cost shape | Runs on | Why |
|
||
|---|---|---|---|
|
||
| Source to camera-linear | per pixel, full res | GPU, the existing pipeline | It *is* the pipeline, stopped early |
|
||
| Keypoint detection | once per frame, at 1024 px | CPU, tract (NEON on the tablet) | Bounded by frame count, not output size. Same runtime faces and masks use. Hand-written WGSL convolutions for a model that runs five times would be work with no visible gain. |
|
||
| Descriptor matching | K² × D per pair | CPU, SIMD | 2048² × 64 × 10 pairs ≈ 3 GFLOP — tens of milliseconds |
|
||
| Rotation solve, bundle adjustment | 3N + 1 parameters, Levenberg–Marquardt | CPU | Microseconds. Not parallel work. |
|
||
| Preview reprojection | per pixel, proxy res | GPU, interactive | Projection and horizon changes re-warp N proxies at frame rate |
|
||
| Full-resolution warp | per output pixel | GPU, chunked (§5) | The heaviest thing in the application |
|
||
| Gain compensation | per overlap region | GPU reduction, then N scalars | Sums, on the histogram pass's pattern (ARCH §5.5) |
|
||
| Seam finding | per overlap pixel | GPU-friendly variant | Graph cut resists the GPU; a distance-transform or per-column DP seam does not. The algorithm is chosen for the GPU, not for the paper. |
|
||
| Multi-band blend | per pixel × levels | GPU, chunked | Laplacian pyramids are separable convolutions — the detail stage's shape |
|
||
| Encode | per pixel, once | CPU, streamed per chunk row | As export does |
|
||
|
||
## 5. Chunked in output space
|
||
|
||
FR-MRG-11 forbids holding the composite as one texture, and two facts force it
|
||
before memory does:
|
||
|
||
- `max_texture_dimension_2d` is 8192 on many mobile GPUs and 16384 on desktop.
|
||
A three-row panorama is routinely 20 000 px wide.
|
||
- Five 24 MP frames at working precision are ~1 GB together. The tablet does
|
||
not have it.
|
||
|
||
**The geometry is known before any full-resolution pixel exists.** Alignment
|
||
runs on proxies; what comes out is a rotation per frame, a focal length, a
|
||
projection and an output rectangle. From those, every output pixel's source
|
||
coordinates in every frame are a closed-form function. That is what makes
|
||
chunking simple rather than clever:
|
||
|
||
```
|
||
for each output chunk C (e.g. 2048 × 2048, in output space):
|
||
frames_in(C) = frames whose projected footprint intersects C
|
||
for each frame F in frames_in(C):
|
||
source tiles T(F, C) = tiles of F that project into C, plus a margin
|
||
render T(F, C) to scene-linear through the pipeline's tile cache
|
||
warp T(F, C) into C's coordinate frame ← GPU
|
||
gain-correct, seam, blend within C ← GPU, with overlap
|
||
read C back, encode its rows ← CPU, streamed
|
||
```
|
||
|
||
The working set is one chunk, its per-frame warped copies, and the source
|
||
tiles that fed them. It does not grow with the composite.
|
||
|
||
**The blend needs a margin.** A Laplacian pyramid of L levels reads
|
||
2^L pixels beyond the chunk edge; a chunk is therefore rendered with a margin
|
||
of that width and the margin discarded after the blend. Seams cross chunk
|
||
boundaries and must agree on both sides: the seam is found once at a reduced
|
||
resolution over the whole overlap (which fits — it is a mask, not an image),
|
||
then upsampled into each chunk. The same is true of gain: the scalars are
|
||
solved once from proxy-resolution overlaps and applied everywhere.
|
||
|
||
**Source tiles are the pipeline's tiles.** ARCH §5.3's cache keys by
|
||
`(VersionId, tile, zoom, graph_hash_prefix)`; the merge asks for tiles of a
|
||
neutral graph at zoom 1 and gets the same caching every other consumer does.
|
||
A tile pulled for one chunk is usually needed by the neighbouring chunk, and
|
||
stays hot for it.
|
||
|
||
### 5.1 The tap — S15.3, answered by reading the composer
|
||
|
||
The fused shader's order, fixed by `operation.rs`'s own tests: warp → as-shot
|
||
white balance → operations → base curve → camera matrix → store. The store is
|
||
either the display encode or, in `OutputMode::LinearWorking`, an unclipped
|
||
`rgba16float` of linear sRGB. That mode exists for the detail stage and is
|
||
selected from the operations, never by a caller flag, so that a shader and
|
||
the texture bound to it cannot disagree.
|
||
|
||
The merge wants the values *before* the curve and matrix (FR-MRG-2), and the
|
||
composer already makes that a matter of uniforms rather than structure: the
|
||
white balance, the matrix and the curve's active flag are all in the reserved
|
||
uniform block, and a fused pass with no operations, `as_shot_wb = 1`,
|
||
`cam_to_srgb = I` and `base_curve_last.z = 0` stores exactly camera-linear
|
||
RGB after the warp. So the tap is:
|
||
|
||
- `EditGraph::compose_camera_linear()` — the `LinearWorking` tail with an
|
||
empty operation list and identity framing, paired by name with
|
||
- `AdjustPass::render_camera_linear()` — binds the f16 target, fills the
|
||
reserved uniforms neutral instead of from the source, returns the texture,
|
||
- and a float readback beside the existing 8-bit one.
|
||
|
||
Nothing in the chain moves. **Precision:** the tap and every chunk buffer
|
||
after it should be `rgba32float`, not f16. A 14-bit sensor has 16 384 steps
|
||
to white; f16 has 2 048 in the top octave, and a composite that is going to
|
||
be re-developed deserves the sensor's precision. The cost is 2× on buffers
|
||
FR-MRG-11 already bounds.
|
||
|
||
**What the DNG carries as a consequence:** the first source's `Make`,
|
||
`Model` and `UniqueCameraModel` — so `base_curve::for_body` finds the 6D's
|
||
curve — its `ColorMatrix1`/`2` with illuminants, and its `AsShotNeutral`. The
|
||
composite then develops through the same profile as its sources, applied
|
||
once. The spike's 64 × 48 file (§8) already carries the matrix and neutral;
|
||
the body name is a string.
|
||
|
||
## 6. The keypoint model
|
||
|
||
FR-MRG-8: works without weights, better with them. The licence read comes
|
||
first (D13's lesson, S15.2).
|
||
|
||
| Model | Licence | Fits tract? | Position |
|
||
|---|---|---|---|
|
||
| **XFeat** (CVPR 2024) | Apache-2.0 | Plain convolutions, fully convolutional, the repo ships an ONNX export | **Chosen.** Fixed 1024 px input, dense heatmap and descriptor map out, NMS and top-K in Rust — the yolo26 pattern |
|
||
| DISK | Apache-2.0 | U-Net, static | Second choice; stronger descriptors, ~3–4× the compute |
|
||
| ALIKE | BSD-3 | Plain convolutions | Fallback if XFeat's export fails F6 |
|
||
| ALIKED | BSD-3 | Deformable convolution in the descriptor head | Unlikely to load |
|
||
| SuperPoint, SuperGlue, R2D2, SiLK, MASt3R | non-commercial | — | Out on licence |
|
||
| LightGlue | Apache-2.0 | Transformer over a variable keypoint count | Not until mutual-nearest-neighbour matching fails on a real set |
|
||
|
||
**S15.2, 2026-09-19: XFeat loads under tract.** `tools/export-xfeat.sh`
|
||
exports the network alone at 768×1024 — thirteen operator types, all
|
||
standard: `Conv`, `InstanceNormalization`, `AveragePool`, `Resize`, `Slice`,
|
||
`Transpose`, `Reshape`, `Concat`, `Add`, `Relu`, `Sigmoid`, `ReduceMean`,
|
||
`Unsqueeze` — and
|
||
[`examples/onnx_probe.rs`](../../core/dr-segment/examples/onnx_probe.rs) loads
|
||
the 2.8 MB file through the app's own `ort`-over-tract backend with nothing
|
||
unsupported, in 28 ms, and runs it in **~300 ms on the reference desktop's
|
||
CPU**. The weights ship as `models/keypoints/xfeat-1024.onnx`, recorded in
|
||
`models/LICENCE.md`. **On the tablet** (S15.4's CPU half, same day):
|
||
`tools/onnx-probe-on-device.sh` cross-builds the probe, and the same file
|
||
runs in **~400 ms per frame** on the reference tablet's NEON cores (ROD2-W09,
|
||
SM8635), with output ranges identical to the desktop's — inside NFR-MRG-1's
|
||
1 s per frame with room to spare, and 1.3× the desktop rather than the 2×
|
||
faces.md §9 measured for its scan. Still to do: a keypoint-level comparison
|
||
against the PyTorch reference once the Rust decoder exists — the probe
|
||
proves the graph runs, not that the numbers match.
|
||
|
||
The outputs are three maps at 1/8 resolution, 96×128 for the export size:
|
||
64-channel descriptors, 65-channel keypoint logits (each 8×8 cell's position
|
||
plus "none"), and a reliability heatmap. The Rust decoder is: softmax over the
|
||
65, pixel-shuffle the first 64 to full resolution, 5×5 non-maximum
|
||
suppression, top-k by reliability, bilinear sampling of the descriptor at
|
||
each keypoint, L2 normalise. That is `detectAndCompute` in the reference,
|
||
minus the network.
|
||
|
||
Without weights: AKAZE (BSD, `akaze` from rust-cv), which is adequate on
|
||
well-textured overlaps and worse on sky, repeated structure and exposure
|
||
drift — which is where a learned detector earns its place.
|
||
|
||
Matching is mutual nearest neighbour with a ratio test, then RANSAC on a
|
||
rotation model. For a panorama — one lens, near-pure rotation, 20–40 %
|
||
overlap — that is what Hugin and OpenCV's stitcher use, and it is enough.
|
||
|
||
## 7. What is ported from where
|
||
|
||
Nothing is linked; everything is read.
|
||
|
||
| Source | Licence | Taken |
|
||
|---|---|---|
|
||
| OpenCV `modules/stitching` | Apache-2.0 | The stage layout — Brown & Lowe (2007) as a set of small classes with one job each — and the warpers' projection maths |
|
||
| OpenPano (ppwwyyxx) | MIT (verify on read) | The estimation and bundle-adjustment maths, function by function, with outputs diffed against it |
|
||
| enblend-enfuse | GPLv2+ | Seam-line optimisation and Burt–Adelson multi-band blending |
|
||
| Hugin `nona` | GPLv2+ | The GLSL remapper, as the reference for the WGSL warp |
|
||
|
||
The golden set (§8 of the requirements) is OpenCV's stitcher on the same
|
||
inputs: a reference output to compare against, within a tolerance calibrated
|
||
the way S9 calibrates R1.
|
||
|
||
## 8. The output file
|
||
|
||
FR-MRG-3. A linear DNG at the source's native scale: `u16` samples on the
|
||
first source's black-subtracted scale, `WhiteLevel` = its white minus its
|
||
black (13 023 for the 6D set: 15 070 − 2 047), `BlackLevel` = 0. Not rescaled
|
||
to 65 535 — the sensor had 14 bits and the file says so, and a value the
|
||
sensor could not have produced is not invented by a multiply. The first
|
||
source's `Make`, `Model`, `UniqueCameraModel`, `ColorMatrix1/2`,
|
||
`CalibrationIlluminant1/2`, `AsShotNeutral` and EXIF are carried, so the
|
||
composite develops through the same profile as its sources. Named from the
|
||
first source with a `-pano` suffix, beside it.
|
||
|
||
Three samples per pixel rather than a CFA: the warp resamples, and there is no
|
||
sensor grid to mosaic back onto. Nothing else about being a RAW is lost —
|
||
no white balance, no curve, no matrix, no clip has been applied — and the
|
||
photographer develops the panorama afterwards as one photograph.
|
||
|
||
The sources are portrait frames in the 6D set: `Orientation` is applied
|
||
before alignment (learned features are not rotation-invariant) and the
|
||
composite is written upright with `Orientation = 1`.
|
||
|
||
Two containers were candidates and S15.1 decided, on 2026-09-19:
|
||
|
||
- **Linear DNG.** `PhotometricInterpretation = LinearRaw`, three samples per
|
||
pixel, `ColorMatrix1` carried from the first source. Re-enters through
|
||
rawler as `Format::Dng` with no new decode path, *if* rawler reads it back.
|
||
What Lightroom writes.
|
||
- **Float TIFF.** `SampleFormat = IEEEFP`, 16 or 32 bits, an ICC profile for
|
||
the working space. Needs `Format::Tiff` and a decode path, but the writer is
|
||
the existing encoder with a different sample type, and nothing about it is
|
||
uncertain.
|
||
|
||
**Linear DNG.** [`examples/linear_dng.rs`](../../core/dr-decode/examples/linear_dng.rs)
|
||
hand-rolls a 64 × 48 `LinearRaw` DNG — one IFD, uncompressed 16-bit RGB,
|
||
`DNGVersion`, `ColorMatrix1`, `AsShotNeutral`, `CalibrationIlluminant1` — and
|
||
rawler 0.7 reads it back: `cpp 3`, the samples interleaved as written, the
|
||
matrix parsed into the camera definition, and `CameraProfile::extract` builds
|
||
the same profile it would for a camera file. ImageMagick's libraw reads the
|
||
same bytes. What does *not* yet work is `dr_decode::decode`, which accepts the
|
||
file as CFA and hands the pipeline three times the samples it expects: the
|
||
`cpp == 3` branch is the work, and it is the only decode work.
|
||
|
||
The composite therefore enters the pipeline as a non-CFA, *linear* source —
|
||
`from_rgba8`'s sibling with `non_linear = false` and the colour matrix carried
|
||
from the DNG — and is developed as any RAW is. The writer is the example's
|
||
IFD, grown up: tiled rather than one strip (FR-MRG-11 encodes per chunk), and
|
||
carrying the first source's EXIF in a sub-IFD as `dr-export` already does.
|
||
|
||
## 9. Interaction
|
||
|
||
- The entry is the grid's selection: two or more images, one action, "Merge
|
||
to panorama". One image, or images from different roots, and the action
|
||
says why it is unavailable.
|
||
- The dialog shows the aligned proxies in the chosen projection, with the
|
||
projection, horizon and crop controls of FR-MRG-4, and the per-frame
|
||
residuals. A frame that failed to align is named there (FR-MRG-5), and the
|
||
merge cannot be confirmed with it in the set.
|
||
- Confirm starts the FR-MRG-7 job. The composite appears in the grid when the
|
||
file is written and catalogued, beside its sources, with the merge as the
|
||
first entry in its history.
|
||
|
||
## 10. Order of work
|
||
|
||
1. **S15**, all four, before anything else. (1) and (2) are a day each and
|
||
either can change the design.
|
||
2. `dr-pano`: keypoints (AKAZE first, XFeat when S15.2 passes), matching,
|
||
RANSAC, rotation solve. Unit-tested against synthetic rotations of one
|
||
frame, where the answer is known exactly.
|
||
3. The working-space tap, and the preview reprojection pass. At this point the
|
||
dialog can show an alignment.
|
||
4. The chunked driver with a feathered blend — the whole path end to end,
|
||
writing a file, before the blend is good.
|
||
5. Gain, seams, multi-band.
|
||
6. The container, the catalog entry, provenance, the history entry.
|
||
7. Tablet: NFR-MRG-1's figure, and FR-MRG-9's ceiling.
|
||
|
||
## 11. Where it stands — 2026-09-19, end of the first day
|
||
|
||
Built, on branch `merge/panorama`, in the order §10 gave:
|
||
|
||
| Piece | Where | State |
|
||
|---|---|---|
|
||
| Geometry: keypoints, matching, homography, focal, bundle adjustment, projections | `core/dr-pano` | Done; 33 tests without a model; the fixture aligns in 4.5 s |
|
||
| XFeat at two shapes under tract | `models/keypoints`, `dr_pano::xfeat` | Done; 300 ms/frame desktop, 400 ms tablet |
|
||
| The camera-space tap | `OutputMode::CameraLinear`, `AdjustPass::render_camera_linear` | Done, `rgba32float`, tiles by view rect |
|
||
| Linear DNG writer, streamed | `dr_export::write_linear_dng` | Done; rawler reads it back |
|
||
| A three-sample `RawImage` re-entering the pipeline | `dr-decode`, `DemosaicedImage::from_linear_rgb16` | Done |
|
||
| Warp, accumulate, resolve, chunk by chunk | `dr_gpu::MergePass`, `merge.wgsl` | Done; feathered blend, scalar gain |
|
||
| The job: load, proxies, align, gains, confirm, merge, provenance | `dr_ui::merge` | Done; `examples/merge.rs` drives it headless |
|
||
| The page: table, preview, projection, Merge/Stop/Back; the grid's button | `merge.slint`, `merge_ui.rs` | Done; `DARKROOM_START_MERGE=a.CR2,b.CR2` lands on it |
|
||
| Placement beside the sources through the outbox, rescan | `merge_ui.rs` | Done, untested against a server |
|
||
|
||
**Measured on the fixture (desktop, 12 × 20 MP, Intel adapter):** proxies
|
||
and keypoints 4 s, alignment 4.5–12.6 s (load-sensitive: the matcher is
|
||
every core), the merge **26 s for a 22 993 × 5 980 composite** in twelve
|
||
bands of 2048 × 512 chunks, 45 s all told, an 825 MB DNG. NFR-MRG-1's 60 s
|
||
holds on the desktop with room; the tablet's figure is still S15.4's open
|
||
half.
|
||
|
||
**Open, in the order they matter:**
|
||
|
||
1. **Auto-crop (FR-MRG-4).** The merge returns a coverage mask per band and
|
||
the file carries the black border. The largest inscribed rectangle over
|
||
the coverage, then the DNG's `DefaultCropOrigin`/`DefaultCropSize`, so
|
||
nothing is thrown away and the develop view opens on the picture.
|
||
2. **Seams and the pyramid** (§10 step 5). The feather hides exposure and
|
||
small misalignment; parallax on the near slope will show as a soft
|
||
double edge at 1:1.
|
||
3. **Vignetting in the tap.** The lens profile's distortion is applied
|
||
before the fetch; its vignetting is an operation and is not. Frame edges
|
||
are darker than their centres by the lens's falloff, and the feather
|
||
averages them into the overlaps.
|
||
4. **The tablet:** memory (twelve 40 MB sensor buffers on the CPU, one
|
||
demosaiced frame at a time on the GPU), the figure, and FR-MRG-9's
|
||
ceiling.
|
||
5. **Horizon and drag-to-correct (FR-MRG-4, the proposed 4a).** The
|
||
alignment failed on nothing in the fixture; the interaction waits for a
|
||
set it fails on.
|
||
6. **`derived_from` names sources by file name**, not content hash: the
|
||
catalog's `content_hash` is null for most images most of the time. The
|
||
hash can join it when the catalog has one.
|
||
|
||
## 12. Filling the border instead of cropping it — MI-GAN, read and measured 2026-09-19
|
||
|
||
Raised after the first merges: the ragged border a cylinder leaves could be
|
||
*filled* rather than cropped away. FR-MRG-4 says no boundary fill, on
|
||
§1.3's "not a pixel editor"; this is the evidence for deciding whether to
|
||
revise that, not a revision.
|
||
|
||
**The candidate: MI-GAN** (Sargsyan et al., ICCV 2023, Picsart AI Research).
|
||
Image inpainting designed for mobile: ~6 M parameters, plain convolutions —
|
||
no FFT, no attention — so it quantises to int8 and runs on a phone's DSP,
|
||
with quality close to LaMa and CoModGAN.
|
||
|
||
**Licence: MIT, code and weights alike** (`LICENSE` and `LICENSE-WEIGHTS`
|
||
in the repository, read the same day). The cleanest position of any model
|
||
in the tree — GPL-compatible, store-compatible, no grant to read around.
|
||
|
||
**Export.** The HuggingFace ONNX files are the *pipeline* — uint8 image and
|
||
mask in, crop-around-mask, resize and blend inside the graph, every
|
||
dimension dynamic — and tract refuses them (F6 again). The bare generator
|
||
exports cleanly from the `migan_512_places2.pt` state dict at a fixed
|
||
`1×4×512×512` (`export_migan.py` in the spike directory; the input is
|
||
`mask − 0.5` and the masked RGB in −1..1, the output RGB in −1..1, the
|
||
caller composites). After slimming the graph is **six operator types**:
|
||
`Add, Clip, Conv, LeakyRelu, Mul, Resize`. 28 MB.
|
||
|
||
**Under tract on the reference desktop: loads in 53 ms, runs in 7.4 s per
|
||
512 × 512 tile, f32.** That is the number. The fixture's border is two
|
||
ragged bands across 22 993 px — roughly ninety 512-px tiles at full
|
||
resolution — so a CPU-f32 fill is ten minutes on the desktop and longer on
|
||
the tablet. Three ways to make it viable, none built:
|
||
|
||
1. **Fill at a quarter of the resolution and upsample.** Sky and scree
|
||
tolerate it; twenty-odd tiles, about three minutes on the desktop CPU. A
|
||
background job with the outbox's patience, not an interactive one.
|
||
2. **int8 on the tablet's Hexagon through QNN**, where the plain-conv design
|
||
is the point and the whole graph should run in milliseconds. The setup
|
||
exists from the eye-state work; MI-GAN is a candidate for the same path.
|
||
3. **A WGSL runtime for those six operators.** A project of its own, and
|
||
the only route that would make it interactive on the desktop.
|
||
|
||
Whichever, the fill is a *proposal* under FR-MRG-1's rule — shown, then
|
||
confirmed — and it would sit beside the crop, not replace it: the crop is
|
||
free and honest, the fill is invented pixels, and the photographer chooses.
|
||
|
||
## 13. The fill, built — 2026-09-19, evening
|
||
|
||
Built the same day on `merge/fill`, on the engine (S16) rather than tract,
|
||
and FR-MRG-4 revised to admit it: the border is *cropped or filled*, the
|
||
photographer's choice, the crop the default.
|
||
|
||
**What runs.** `dr_pano::fill` is the engine-independent half: an
|
||
`Inpainter` trait (a 512-px tile in, the same tile out) and `fill_border`,
|
||
which owns everything the model does not — which tiles, what context, how
|
||
to blend. `dr_pano::migan::MiGan` is the trait over the shipped generator
|
||
under `dr_inference_engine` with the new `Role::Inpainter`, so it takes
|
||
whichever rung the device has. The merge job runs the fill at **half the
|
||
composite's resolution**, in a display-ish space (white balance, camera
|
||
matrix, gamma — invertible, so the result goes back to camera-linear and
|
||
into the same linear DNG), and the full-resolution merge samples the fill
|
||
where no frame reached.
|
||
|
||
**What the spike taught, tried in order and kept or dropped.**
|
||
|
||
1. *Context across the coverage edge.* MI-GAN was trained on holes inside
|
||
pictures; given a hole at the picture's edge it invents a structure along
|
||
the open side (white streaks in the sky, on the first try). The known
|
||
content is therefore **mirrored** across the coverage edge into the hole
|
||
and into a 256-px ring, column by column for the top and bottom bands
|
||
and row by row for the sides; the model interpolates between real and
|
||
mirrored sky rather than extrapolating into nothing. *Replicated* rows
|
||
(the edge row continued flat) streaked the grass; a detrended mix (tone
|
||
replicated, texture mirrored) smeared; a low-pass extrapolation banded.
|
||
Mirror stays.
|
||
2. *Coarse to fine.* One pass at the working resolution let the boundary
|
||
leak in — each 512 tile saw only its own corner of the hole. So a
|
||
**coarse pass at a quarter** decides the structure with the whole border
|
||
in a few tiles, and **fine passes in 96-px bands** from the real edge
|
||
outward regenerate texture, each band the only unknown with the previous
|
||
band on its near side and the upsampled coarse fill on its far side.
|
||
3. *The seam.* A hard cut between real and invented showed as a sharpness
|
||
step. The known mask is eroded by a **24-px feather** (48 at half
|
||
resolution) and the fill blended in across that margin by distance to
|
||
the real edge, smoothstep.
|
||
4. *Partial pixels.* The seams were still visible until the cause was found
|
||
upstream of the fill: the camera-space tap stored **black with alpha 1**
|
||
for a pixel the lens correction pushed off the sensor, and the warp
|
||
averaged it in — a dark, poorly interpolated fringe along every frame's
|
||
edge that the fill then continued. `OutputMode::CameraLinear` now
|
||
stores alpha 0 for a pixel that is not there and the merge's warp
|
||
weights by the sampled alpha, so the fringe never enters the composite.
|
||
The mask erosion before the fill dropped from 16 px to 4.
|
||
|
||
5. *What is still wrong, and why it ships anyway.* With the seams gone the
|
||
content itself is the problem in the deep corners: the model, trained
|
||
on Places2, puts bright cloud-and-peak shapes into a sky hole and a
|
||
water-like band under grass — its prior for "top of a picture" and
|
||
"bottom of a landscape", not anything in the context (the same shapes
|
||
appear with the mirror capped, uncapped, and on the CPU as on TensorRT).
|
||
Thin borders are fine; that is most of a hand-held sweep. So the fill
|
||
ships **experimental**: opt-in, previewed, its knobs on the page and
|
||
in the sidecar, and `cargo run -p dr-ui --example fill` re-runs any
|
||
merge's dumped input (`DR_FILL_DUMP=dir`) stage by stage in seconds so
|
||
the next attempt is made from the picture, not from a seven-minute
|
||
merge. Candidates for that attempt: a context that is not a mirror at
|
||
all in deep holes (the coarse pass's own answer, iterated), a sky
|
||
detector that fills sky by extrapolating the gradient and leaves the
|
||
model to texture, or a different model.
|
||
|
||
**Measured, the fixture's twelve frames (22 991 × 5 978), 348 tiles at
|
||
half resolution.** 312 s on ONNX Runtime's CPU pool on the reference
|
||
desktop (≈ 0.8 s a tile). On TensorRT fp16: **100 s**, of which 60 ms a
|
||
tile was the engine hashing the 28 MB model on every acquire (fixed, the
|
||
hash is taken at open) and 150 ms a tile the GPU itself — throttled:
|
||
`trtexec` on the same engine read 23 ms at noon on a cool machine and
|
||
152 ms that evening after two hours of builds, nvidia-smi showing SW power
|
||
cap and thermal slowdown. Cool, the fill is ~10 s. The TensorRT engine
|
||
compiles once, in 13 minutes, cached under the inference directory.
|
||
|
||
**The runtime is a packaging matter.** Arch's `onnxruntime-opt-cuda` has
|
||
no TensorRT provider ("not enabled in this build") and its CUDA provider
|
||
does not load against cuDNN 9, so on this machine the app fell to ORT CPU
|
||
until the official `onnxruntime-linux-x64-gpu_cuda13` tarball (1.30.0,
|
||
which links the system CUDA 13.4 and TensorRT 10.16) was unpacked and
|
||
named with `DARKROOM_ORT_DIR`; `/usr/lib/darkroom` is searched too, for a
|
||
package that ships it. §12's point 2 for the tablet is unchanged.
|
||
|
||
**On the page.** A *Border* choice beside the projection — *Crop to the
|
||
picture* / *Fill the border* — with a caption saying what the fill is; the
|
||
preview re-renders filled when chosen, at preview resolution, so the
|
||
choice is seen before it is confirmed (FR-MRG-1). Greyed out with the reason
|
||
when `migan-512.onnx` is not in the model directory. Under the fill, while
|
||
it is experimental, its six knobs as sliders — working scale, edge
|
||
erosion, coarse pass, band width, mirror depth, seam feather — each
|
||
committing a redraw of the preview. A filled merge's sidecar says `border
|
||
filled` with the knobs used, and its default crop is still the inscribed
|
||
rectangle.
|
||
|
||
**Ships.** `models/inpaint/migan-512.onnx` (LFS, 28 MB, MIT,
|
||
`models/LICENCE.md`), installed by the PKGBUILD and unpacked by the APK
|
||
beside the face and scene models; `tools/export-migan.sh` regenerates it
|
||
from the upstream checkpoint.
|
||
|
||
## 14. The fill, trained — 2026-09-20
|
||
|
||
§13.5 named the remaining fault: in a deep corner the stock model puts its
|
||
Places2 prior — clouds, peaks, a water line — into a hole, because it was
|
||
trained on holes *inside* pictures and a panorama's border is a hole with
|
||
the picture on one side and nothing on the other. Every ring (mirror,
|
||
replicate, detrend) treated the symptom. The fix is a model that has seen
|
||
the real thing: **MI-GAN's 512 generator fine-tuned on border-shaped voids
|
||
cut from the user's own photographs**, in a separate repository
|
||
(`darkroom-infill`, beside this one), so the truth beyond the void is known
|
||
and the model learns one-sided extrapolation.
|
||
|
||
**What it was trained on.** Voids made the way this merge makes them:
|
||
frames with a yaw, a common pitch and per-frame roll, projected onto the
|
||
cylinder and rasterised, the canvas their union's bounding box, the void
|
||
the canvas outside the union — arcs where straight edges bent, cusps where
|
||
frames meet, the bow-tie wedge at a corner (a third of tiles are cut at a
|
||
canvas corner). Voids to 256 px deep at a 512 tile. Half the deep tiles
|
||
train the *second pass*: a no-grad first pass fills the tile, its nearest
|
||
band (64–256 px) is marked known, and the remaining void is the example —
|
||
so the model continues its own output without drift, which is how `fill`
|
||
runs deep voids. Data: ~7 400 pictures — the 1024-px proxy tier of the
|
||
library and ~2 000 raws sampled evenly across every year, developed at
|
||
half size. Loss: hole-weighted L1, VGG16 perceptual, a hinge PatchGAN.
|
||
One night on the reference desktop's RTX 3050.
|
||
|
||
**What changed here.** `FillParams::mirror_depth` **0** is now "no ring":
|
||
the void reaches the tile's edge with nothing beyond, and beyond the band
|
||
being filled the void stays *unknown* rather than presented as known coarse
|
||
fill — the two conditions the model was trained under. Defaults: mirror 0,
|
||
coarse 1 (the coarse pass seeds nothing the model is allowed to see), band
|
||
192. The ring remains on the page for the stock model's sake, at any depth
|
||
above zero. The model file is a drop-in (`models/inpaint/migan-512.onnx`,
|
||
same six operators, same tensors) and the engine loads it unchanged.
|
||
|
||
**Measured, 240 held-out tiles with projection-shaped voids (PSNR in the
|
||
hole, dB / LPIPS on the composite), stock → shipped (step 3 607):** edge
|
||
16.9 → 18.5 / 0.121 → 0.136; corner 14.6 → 16.3 / 0.183 → 0.205; interior
|
||
18.2 → 19.3 / 0.051 → 0.056. Read both columns: the fine-tune gains ~2 dB
|
||
on edges and corners because it stops inventing objects, and *loses* on
|
||
LPIPS because what it paints in a deep void is smoother than the stock
|
||
model's confident wrong texture — LPIPS rewards texture, right or wrong.
|
||
On the fixture's dump at half resolution (the merge's working size) the
|
||
sky corners are sky, with no structure and a faint tone step at worst;
|
||
the ground bands carry a fine texture at the right tone, softer than the
|
||
real scree above them. The stock model's top-left corner on the same
|
||
dump is a glowing invented structure. The training's own record — what
|
||
each loss weighting did, and the two runs abandoned (blur under L1 in
|
||
the hole; a brick pattern under a strong adversarial term against a
|
||
discriminator that had not learned) — is `runs/` in `darkroom-infill`.
|
||
|
||
**What is still wrong.** The ground fill is softer than its context —
|
||
texture, not structure, is what a night on a laptop GPU could not finish.
|
||
The levers, in order: a discriminator that learns (a pretrained one —
|
||
MI-GAN's own from the unfused checkpoint — instead of a PatchGAN from
|
||
scratch), feature matching, and more steps at 512. FR-MRG-4's
|
||
*experimental* stays.
|