From d04b8f604474badf02318d9f0affbb35fdd01e66 Mon Sep 17 00:00:00 2001 From: Duncan Tourolle Date: Sat, 19 Sep 2026 20:34:14 +0200 Subject: [PATCH] =?UTF-8?q?FR-MRG-4:=20the=20border=20is=20cropped=20or=20?= =?UTF-8?q?filled,=20the=20fill=20experimental;=20panorama.md=20=C2=A713?= =?UTF-8?q?=20records=20what=20was=20built=20and=20measured?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- docs/panorama.md | 98 ++++++++++++++++++++++++++++++++++++++++++++ docs/requirements.md | 25 ++++++++--- 2 files changed, 117 insertions(+), 6 deletions(-) diff --git a/docs/panorama.md b/docs/panorama.md index b84b784..be2d80c 100644 --- a/docs/panorama.md +++ b/docs/panorama.md @@ -388,3 +388,101 @@ the tablet. Three ways to make it viable, none built: Whichever, the fill is a *proposal* under FR-MRG-1's rule — shown, then confirmed — and it would sit beside the crop, not replace it: the crop is free and honest, the fill is invented pixels, and the photographer chooses. + +## 13. The fill, built — 2026-09-19, evening + +Built the same day on `merge/fill`, on the engine (S16) rather than tract, +and FR-MRG-4 revised to admit it: the border is *cropped or filled*, the +photographer's choice, the crop the default. + +**What runs.** `dr_pano::fill` is the engine-independent half: an +`Inpainter` trait (a 512-px tile in, the same tile out) and `fill_border`, +which owns everything the model does not — which tiles, what context, how +to blend. `dr_pano::migan::MiGan` is the trait over the shipped generator +under `dr_inference_engine` with the new `Role::Inpainter`, so it takes +whichever rung the device has. The merge job runs the fill at **half the +composite's resolution**, in a display-ish space (white balance, camera +matrix, gamma — invertible, so the result goes back to camera-linear and +into the same linear DNG), and the full-resolution merge samples the fill +where no frame reached. + +**What the spike taught, tried in order and kept or dropped.** + +1. *Context across the coverage edge.* MI-GAN was trained on holes inside + pictures; given a hole at the picture's edge it invents a structure along + the open side (white streaks in the sky, on the first try). The known + content is therefore **mirrored** across the coverage edge into the hole + and into a 256-px ring, column by column for the top and bottom bands + and row by row for the sides; the model interpolates between real and + mirrored sky rather than extrapolating into nothing. *Replicated* rows + (the edge row continued flat) streaked the grass; a detrended mix (tone + replicated, texture mirrored) smeared; a low-pass extrapolation banded. + Mirror stays. +2. *Coarse to fine.* One pass at the working resolution let the boundary + leak in — each 512 tile saw only its own corner of the hole. So a + **coarse pass at a quarter** decides the structure with the whole border + in a few tiles, and **fine passes in 96-px bands** from the real edge + outward regenerate texture, each band the only unknown with the previous + band on its near side and the upsampled coarse fill on its far side. +3. *The seam.* A hard cut between real and invented showed as a sharpness + step. The known mask is eroded by a **24-px feather** (48 at half + resolution) and the fill blended in across that margin by distance to + the real edge, smoothstep. +4. *Partial pixels.* The seams were still visible until the cause was found + upstream of the fill: the camera-space tap stored **black with alpha 1** + for a pixel the lens correction pushed off the sensor, and the warp + averaged it in — a dark, poorly interpolated fringe along every frame's + edge that the fill then continued. `OutputMode::CameraLinear` now + stores alpha 0 for a pixel that is not there and the merge's warp + weights by the sampled alpha, so the fringe never enters the composite. + The mask erosion before the fill dropped from 16 px to 4. + +5. *What is still wrong, and why it ships anyway.* With the seams gone the + content itself is the problem in the deep corners: the model, trained + on Places2, puts bright cloud-and-peak shapes into a sky hole and a + water-like band under grass — its prior for "top of a picture" and + "bottom of a landscape", not anything in the context (the same shapes + appear with the mirror capped, uncapped, and on the CPU as on TensorRT). + Thin borders are fine; that is most of a hand-held sweep. So the fill + ships **experimental**: opt-in, previewed, its knobs on the page and + in the sidecar, and `cargo run -p dr-ui --example fill` re-runs any + merge's dumped input (`DR_FILL_DUMP=dir`) stage by stage in seconds so + the next attempt is made from the picture, not from a seven-minute + merge. Candidates for that attempt: a context that is not a mirror at + all in deep holes (the coarse pass's own answer, iterated), a sky + detector that fills sky by extrapolating the gradient and leaves the + model to texture, or a different model. + +**Measured, the fixture's twelve frames (22 991 × 5 978), 348 tiles at +half resolution.** 312 s on ONNX Runtime's CPU pool on the reference +desktop (≈ 0.8 s a tile). On TensorRT fp16: **100 s**, of which 60 ms a +tile was the engine hashing the 28 MB model on every acquire (fixed, the +hash is taken at open) and 150 ms a tile the GPU itself — throttled: +`trtexec` on the same engine read 23 ms at noon on a cool machine and +152 ms that evening after two hours of builds, nvidia-smi showing SW power +cap and thermal slowdown. Cool, the fill is ~10 s. The TensorRT engine +compiles once, in 13 minutes, cached under the inference directory. + +**The runtime is a packaging matter.** Arch's `onnxruntime-opt-cuda` has +no TensorRT provider ("not enabled in this build") and its CUDA provider +does not load against cuDNN 9, so on this machine the app fell to ORT CPU +until the official `onnxruntime-linux-x64-gpu_cuda13` tarball (1.30.0, +which links the system CUDA 13.4 and TensorRT 10.16) was unpacked and +named with `DARKROOM_ORT_DIR`; `/usr/lib/darkroom` is searched too, for a +package that ships it. §12's point 2 for the tablet is unchanged. + +**On the page.** A *Border* choice beside the projection — *Crop to the +picture* / *Fill the border* — with a caption saying what the fill is; the +preview re-renders filled when chosen, at preview resolution, so the +choice is seen before it is confirmed (FR-MRG-1). Greyed out with the reason +when `migan-512.onnx` is not in the model directory. Under the fill, while +it is experimental, its six knobs as sliders — working scale, edge +erosion, coarse pass, band width, mirror depth, seam feather — each +committing a redraw of the preview. A filled merge's sidecar says `border +filled` with the knobs used, and its default crop is still the inscribed +rectangle. + +**Ships.** `models/inpaint/migan-512.onnx` (LFS, 28 MB, MIT, +`models/LICENCE.md`), installed by the PKGBUILD and unpacked by the APK +beside the face and scene models; `tools/export-migan.sh` regenerates it +from the upstream checkpoint. diff --git a/docs/requirements.md b/docs/requirements.md index 7e97af8..8075812 100644 --- a/docs/requirements.md +++ b/docs/requirements.md @@ -1729,15 +1729,28 @@ profile, exposure, everything in §3.3 — as one photograph, from the sensor's **FR-MRG-4 — Projection and framing.** Cylindrical, spherical or perspective, chosen from the field of view and overridable; the horizon levelled from the estimated rotations, overridable by a -drag; auto-crop to the largest inscribed rectangle, overridable. No boundary fill: painting pixels -that were never captured is the pixel editing §1.3 excludes. +drag; the border either cropped to the largest inscribed rectangle or filled, the photographer's +choice, the crop the default. *The auto-crop is non-destructive* (built 2026-09-19): it is the composite's default crop, not a cut — the whole merge including its border is in the file, and resetting the crop shows it. -*Boundary fill is an open question, raised the same day:* a fill of the border from the picture's -own edge through FR-DEV-8's heal would be non-generative and honest about what it is; a generative -inpainter is a model licence, hundreds of megabytes of weights and an inference path the -application does not have. Neither is decided; the clause stands as written until one is. + +*The fill is generative and opt-in* (revised the same day, having first said "no boundary +fill"): MI-GAN (Sargsyan et al., ICCV 2023; MIT code and weights, `models/LICENCE.md`) paints +the uncovered border from the picture's own edge, under the inference engine (S16, [inference.md](inference.md)). It is a +proposal under FR-MRG-1 — shown on the page, chosen against the crop, confirmed before the merge +— never the default, and a merge that used it says so in its sidecar (`border filled`) so the +invented pixels are declared, not passed off as captured. §1.3 still holds: the fill touches only +pixels no frame reached, never the photograph; a filled merge keeps the inscribed crop as its +default so the honest picture is one reset away. The filler is loaded from the model directory +when present and the choice is greyed out, with the reason, when it is not. + +*Experimental, as shipped 2026-09-19.* The fill is right in thin borders and wrong in deep +corners, where the model invents cloud and water where there is sky and grass +(panorama.md §13); it ships opt-in with **every knob on the page** — working scale, edge +erosion, coarse pass, band width, mirror depth, seam feather — each redrawing the preview, and +the knobs used are written into the sidecar's `merge` line beside `border filled`. The knobs +leave the page when the defaults are right; the sidecar record stays. **FR-MRG-5 — Honesty of failure.** *(general to any merge)* A frame that cannot be aligned is named, with why — too few matches, no overlap with any other frame, a residual above the stated