Specify panorama merging: §3.11, D18, S15, and the design in panorama.md

A merge writes a new source file beside its sources (D18) rather than a
multi-source Version, which answers the schema question §7 had been holding
open for panorama, HDR merge and focus stacking together. The panorama is
undeferred as FR-MRG-1 … 11; the other two stay in §7 with their data model
decided.

FR-MRG-10 and 11 fix where the work runs — every per-pixel stage on the GPU,
the composite never held as one texture — because the output exceeds
max_texture_dimension_2d before it exceeds memory. panorama.md carries the
stage table, the chunked output driver, the model licences and the porting
sources. S15 gates all of it.

Coverage falls from 83.0% to 77.2%: thirteen requirements entered with no
code, and outstanding.md §11 says so.
This commit is contained in:
2026-09-19 15:24:10 +02:00
parent f79a76f2d5
commit c901fc1a0a
4 changed files with 470 additions and 84 deletions
+14 -1
View File
@@ -455,7 +455,20 @@ whether something *should* be built — which is the opposite of the order §9 a
--- ---
## 11. D12, which governed all of the above ## 11. Merging — specified 2026-09-19, nothing built
§3.11 was written on 2026-09-19 under D18, undeferring the panorama from §7 and leaving HDR merge
and focus stacking there with their data model decided. Eleven `FR-MRG` clauses and two `NFR-MRG`
figures entered the register at once with no code behind any of them, which is why the coverage
figure fell from 83.0% to 77.2% on the same day — a specification, not a regression.
[panorama.md](panorama.md) is the design, and its §10 is the order of work. Nothing starts before
**S15**: whether rawler reads back a linear DNG the application writes, whether XFeat loads under
tract at a fixed shape, whether the working-space texture can be tapped where FR-MRG-2 needs it,
and what a chunked blend of a 100 MP composite costs on the tablet. The first two are a day each
and either can change the design, which is the reason they come first.
## 12. D12, which governed all of the above
> **Decided 2026-09-19, by events.** The scope stands as calibrated, v1 has no date, and `(post-v1)` > **Decided 2026-09-19, by events.** The scope stands as calibrated, v1 has no date, and `(post-v1)`
> in §7 is the one way a clause leaves the count — used for the plugin API and nothing else. The > in §7 is the one way a clause leaves the count — used for the plugin API and nothing else. The
+211
View File
@@ -0,0 +1,211 @@
# Panorama
**Status:** Draft · 2026-09-19
**Companion to:** [requirements.md](requirements.md) §3.11 FR-MRG-1 … 11, D18, S15 · [architecture.md](architecture.md) §5.2, §6.2
The first merge (§3.11): several frames, rotated about one point, become one
photograph. This document is how that lands on the pipeline that exists now —
which stages, where each runs, how the composite is produced in chunks when it
is larger than any texture or any memory, what is ported from where, and what
the keypoint model may be under D8.
---
## 1. Why it is worth the work
The audience shoots panoramas and leaves the application to stitch them. That
is the same workflow break dust was (`spot-removal.md` §1): a RAW editor that
does everything but the one thing, and the photographer's work ends up in a
JPEG produced by a tool that never saw the RAW.
It is also the merge whose alignment problem is smallest. A panorama is a
rotation — three parameters per frame plus a focal length — with no depth to
recover. HDR merge and focus stacking share its data model (D18) and most of
its machinery (FR-MRG-3, 5, 6, 7, 10, 11 are written to be general); building
the panorama first builds the shared part on the easiest geometry.
## 2. Non-goals
- **Not structure-from-motion.** No translation is solved for. A hand-held set
with parallax gets its ghosts hidden by seam placement, and a set with real
parallax is not a panorama. COLMAP's front end is the right mental model;
its back end is the wrong problem.
- **Not a multi-source Version.** D18. The composite is a file, and nothing in
the catalog, the sidecar format or sync learns about cross-references.
- **Not boundary fill.** Painting pixels that were never captured is the pixel
editing §1.3 excludes. Auto-crop is the tool.
- **Not automatic.** The tool proposes an alignment and writes nothing until
the photographer confirms. Same rule as spot removal and D17, for the same
reason: a merge that silently omits or misplaces a frame is the failure this
application must not have.
- **Not HDR-panorama in one pass.** Until HDR merge exists on its own, a
bracketed panorama is bracketed frames merged first, then stitched.
## 3. What is new, precisely
Nearly all of it, unlike spot removal. The pipeline renders one source to one
texture; nothing in the tree detects keypoints, estimates a rotation, warps
into a projection, finds a seam, or blends a pyramid. What exists and is
reused:
| Exists | Where | Reused for |
|---|---|---|
| Render a source to scene-linear on the GPU | `dr-gpu` demosaic → camera profile → working space | FR-MRG-2's input, once a tap after lens correction and before tone exists (S15.3) |
| Tiled rendering with a priority scheduler | ARCH §5.3 | Pulling source tiles on demand into an output chunk (§5 below) |
| A non-CFA source entering the pipeline | `Demosaicer::from_rgba8` | The composite's decode path, if the container is a TIFF (S15.1) |
| DNG matrices read through rawler | `dr-decode::profile` | The composite's decode path, if the container is a DNG |
| Static-shape ONNX under tract, heads decoded in Rust | `dr-segment` | The keypoint model (§6) |
| A batch worker with its own `GpuContext`, activity row, cancel | `dr-ui::export` | FR-MRG-7 verbatim |
| The 16-bit TIFF encoder with metadata sub-IFDs | `dr-export::encode` | FR-MRG-3's writer, extended to linear samples |
| Multi-select in the grid | `collections_ui::selected` | The entry point |
New: a `core/dr-pano` crate holding the geometry (keypoints, matching, the
rotation solve), a set of WGSL passes in `dr-gpu` (reprojection, gain,
seam, pyramid blend), the chunked output driver, the container writer, and
the dialog.
## 4. The stages, and where each runs
FR-MRG-10 states the rule; this is the table it was written from.
| Stage | Cost shape | Runs on | Why |
|---|---|---|---|
| Source to scene-linear | per pixel, full res | GPU, the existing pipeline | It *is* the pipeline |
| Keypoint detection | once per frame, at 1024 px | CPU, tract (NEON on the tablet) | Bounded by frame count, not output size. Same runtime faces and masks use. Hand-written WGSL convolutions for a model that runs five times would be work with no visible gain. |
| Descriptor matching | K² × D per pair | CPU, SIMD | 2048² × 64 × 10 pairs ≈ 3 GFLOP — tens of milliseconds |
| Rotation solve, bundle adjustment | 3N + 1 parameters, Levenberg–Marquardt | CPU | Microseconds. Not parallel work. |
| Preview reprojection | per pixel, proxy res | GPU, interactive | Projection and horizon changes re-warp N proxies at frame rate |
| Full-resolution warp | per output pixel | GPU, chunked (§5) | The heaviest thing in the application |
| Gain compensation | per overlap region | GPU reduction, then N scalars | Sums, on the histogram pass's pattern (ARCH §5.5) |
| Seam finding | per overlap pixel | GPU-friendly variant | Graph cut resists the GPU; a distance-transform or per-column DP seam does not. The algorithm is chosen for the GPU, not for the paper. |
| Multi-band blend | per pixel × levels | GPU, chunked | Laplacian pyramids are separable convolutions — the detail stage's shape |
| Encode | per pixel, once | CPU, streamed per chunk row | As export does |
## 5. Chunked in output space
FR-MRG-11 forbids holding the composite as one texture, and two facts force it
before memory does:
- `max_texture_dimension_2d` is 8192 on many mobile GPUs and 16384 on desktop.
A three-row panorama is routinely 20 000 px wide.
- Five 24 MP frames at working precision are ~1 GB together. The tablet does
not have it.
**The geometry is known before any full-resolution pixel exists.** Alignment
runs on proxies; what comes out is a rotation per frame, a focal length, a
projection and an output rectangle. From those, every output pixel's source
coordinates in every frame are a closed-form function. That is what makes
chunking simple rather than clever:
```
for each output chunk C (e.g. 2048 × 2048, in output space):
frames_in(C) = frames whose projected footprint intersects C
for each frame F in frames_in(C):
source tiles T(F, C) = tiles of F that project into C, plus a margin
render T(F, C) to scene-linear through the pipeline's tile cache
warp T(F, C) into C's coordinate frame ← GPU
gain-correct, seam, blend within C ← GPU, with overlap
read C back, encode its rows ← CPU, streamed
```
The working set is one chunk, its per-frame warped copies, and the source
tiles that fed them. It does not grow with the composite.
**The blend needs a margin.** A Laplacian pyramid of L levels reads
2^L pixels beyond the chunk edge; a chunk is therefore rendered with a margin
of that width and the margin discarded after the blend. Seams cross chunk
boundaries and must agree on both sides: the seam is found once at a reduced
resolution over the whole overlap (which fits — it is a mask, not an image),
then upsampled into each chunk. The same is true of gain: the scalars are
solved once from proxy-resolution overlaps and applied everywhere.
**Source tiles are the pipeline's tiles.** ARCH §5.3's cache keys by
`(VersionId, tile, zoom, graph_hash_prefix)`; the merge asks for tiles of a
neutral graph at zoom 1 and gets the same caching every other consumer does.
A tile pulled for one chunk is usually needed by the neighbouring chunk, and
stays hot for it.
## 6. The keypoint model
FR-MRG-8: works without weights, better with them. The licence read comes
first (D13's lesson, S15.2).
| Model | Licence | Fits tract? | Position |
|---|---|---|---|
| **XFeat** (CVPR 2024) | Apache-2.0 | Plain convolutions, fully convolutional, the repo ships an ONNX export | **Chosen.** Fixed 1024 px input, dense heatmap and descriptor map out, NMS and top-K in Rust — the yolo26 pattern |
| DISK | Apache-2.0 | U-Net, static | Second choice; stronger descriptors, ~3–4× the compute |
| ALIKE | BSD-3 | Plain convolutions | Fallback if XFeat's export fails F6 |
| ALIKED | BSD-3 | Deformable convolution in the descriptor head | Unlikely to load |
| SuperPoint, SuperGlue, R2D2, SiLK, MASt3R | non-commercial | — | Out on licence |
| LightGlue | Apache-2.0 | Transformer over a variable keypoint count | Not until mutual-nearest-neighbour matching fails on a real set |
Without weights: AKAZE (BSD, `akaze` from rust-cv), which is adequate on
well-textured overlaps and worse on sky, repeated structure and exposure
drift — which is where a learned detector earns its place.
Matching is mutual nearest neighbour with a ratio test, then RANSAC on a
rotation model. For a panorama — one lens, near-pure rotation, 20–40 %
overlap — that is what Hugin and OpenCV's stitcher use, and it is enough.
## 7. What is ported from where
Nothing is linked; everything is read.
| Source | Licence | Taken |
|---|---|---|
| OpenCV `modules/stitching` | Apache-2.0 | The stage layout — Brown & Lowe (2007) as a set of small classes with one job each — and the warpers' projection maths |
| OpenPano (ppwwyyxx) | MIT (verify on read) | The estimation and bundle-adjustment maths, function by function, with outputs diffed against it |
| enblend-enfuse | GPLv2+ | Seam-line optimisation and Burt–Adelson multi-band blending |
| Hugin `nona` | GPLv2+ | The GLSL remapper, as the reference for the WGSL warp |
The golden set (§8 of the requirements) is OpenCV's stitcher on the same
inputs: a reference output to compare against, within a tolerance calibrated
the way S9 calibrates R1.
## 8. The output file
FR-MRG-3. Scene-linear, ≥ 16 bits, wide gamut, the first source's capture
metadata, named from the first source with a `-pano` suffix, beside it.
Two containers are candidates and S15.1 decides:
- **Linear DNG.** `PhotometricInterpretation = LinearRaw`, three samples per
pixel, `ColorMatrix1` carried from the first source. Re-enters through
rawler as `Format::Dng` with no new decode path, *if* rawler reads it back.
What Lightroom writes.
- **Float TIFF.** `SampleFormat = IEEEFP`, 16 or 32 bits, an ICC profile for
the working space. Needs `Format::Tiff` and a decode path, but the writer is
the existing encoder with a different sample type, and nothing about it is
uncertain.
Either way the composite enters the pipeline as a non-CFA, *linear* source —
`from_rgba8`'s sibling with `non_linear = false` and the base curve resolved
from the carried matrix — and is developed as any RAW is.
## 9. Interaction
- The entry is the grid's selection: two or more images, one action, "Merge
to panorama". One image, or images from different roots, and the action
says why it is unavailable.
- The dialog shows the aligned proxies in the chosen projection, with the
projection, horizon and crop controls of FR-MRG-4, and the per-frame
residuals. A frame that failed to align is named there (FR-MRG-5), and the
merge cannot be confirmed with it in the set.
- Confirm starts the FR-MRG-7 job. The composite appears in the grid when the
file is written and catalogued, beside its sources, with the merge as the
first entry in its history.
## 10. Order of work
1. **S15**, all four, before anything else. (1) and (2) are a day each and
either can change the design.
2. `dr-pano`: keypoints (AKAZE first, XFeat when S15.2 passes), matching,
RANSAC, rotation solve. Unit-tested against synthetic rotations of one
frame, where the answer is known exactly.
3. The working-space tap, and the preview reprojection pass. At this point the
dialog can show an alignment.
4. The chunked driver with a feathered blend — the whole path end to end,
writing a file, before the blend is good.
5. Gain, seams, multi-band.
6. The container, the catalog entry, provenance, the history entry.
7. Tablet: NFR-MRG-1's figure, and FR-MRG-9's ceiling.
+151 -2
View File
@@ -1678,6 +1678,114 @@ application starts, opens a photograph, names the responsible plugin, and contin
--- ---
### 3.11 Merging images
Several photographs become one. Panorama is the first merge and the only one specified; HDR merge
and focus stacking share its data model (D18) and are still deferred in §7. The clauses below are
written for the panorama and, where a clause is general to any merge, say so.
**FR-MRG-1 — Panorama from a selection.** Two or more selected images are aligned and blended
into one composite, which is written as a new source file per D18. The tool is never automatic:
it proposes an alignment, the photographer sees it and confirms, and nothing is written before
that press.
The stated audience (D11) shoots panoramas and currently leaves the application to stitch them,
which is the workflow break FR-DEV-8 was added to close for dust. It is also the first of the
three §7 merges, and the one whose alignment problem is smallest — a rotation about one point,
with no depth to recover — so it is where the shared machinery is built.
**FR-MRG-2 — What is stitched.** Each source enters the merge at develop-neutral scene-linear:
after black and white levels, demosaic, camera profile and lens distortion correction, before any
tone or colour adjustment, with one white balance — the first frame's — applied to all. The
sources' own edits are not baked in. The composite is developed afterwards as if it were a new
RAW.
This is the clause that decides what the output *is*. Stitching the rendered edits is what a JPEG
stitcher does; the result cannot be re-developed, and any difference between the frames' edits
becomes a seam. Stitching neutral pixels produces something that behaves like a photograph the
camera could have taken, and every develop operation in §3.3 then applies to it once, not five
times. Lens correction sits above the cut because a distorted frame does not align; white balance
sits above it because the scalars must agree across frames or the overlaps do not match.
**FR-MRG-3 — The output file.** *(general to any merge)* Scene-linear, at least 16 bits per
channel, in a wide gamut with the colour transform resolved or the profile carried, with capture
metadata from the first source. Named from the first source with a stated suffix and placed
beside it. Where the sources' folder is not writable — a remote-only tier, a read-only mount —
it goes where an export goes (FR-EXP-6) and the interface says so before the merge starts.
The container is a spike result (S15), not a requirement: a linear DNG if rawler reads back what
the application writes, else a float TIFF with a decode path of its own. The choice is invisible
to everything above the decoder.
**FR-MRG-4 — Projection and framing.** Cylindrical, spherical or perspective, chosen from the
field of view and overridable; the horizon levelled from the estimated rotations, overridable by a
drag; auto-crop to the largest inscribed rectangle, overridable. No boundary fill: painting pixels
that were never captured is the pixel editing §1.3 excludes.
**FR-MRG-5 — Honesty of failure.** *(general to any merge)* A frame that cannot be aligned is
named, with why — too few matches, no overlap with any other frame, a residual above the stated
bound — and the merge stops. Never a silent drop, never a best-effort composite with a frame
missing.
The same rule as `spot-removal.md`'s and D17's: a tool that quietly alters or omits part of a
photograph is the failure this application must not have, and here the omission would be an
entire frame.
**FR-MRG-6 — Provenance.** *(general to any merge)* The composite's sidecar carries
`derived_from`: the content hashes of its sources in order, and the merge parameters. The history
records the merge as the first entry, and export metadata declares the composite as one. Sources
trashed later leave the list dangling; the panel says so and nothing is blocked.
Provenance, not dependency. The composite renders from itself alone; `derived_from` exists so the
photographer, and anyone they hand the file to, can see what it is made of. The audience is
RAW-literate and a composite declares itself (D17). Nothing here specifies C2PA.
**FR-MRG-7 — Execution.** *(general to any merge)* A background job on the pattern FR-EXP-7
established: its own thread, its own `GpuContext`, a row in the activity panel, cancellable with
NFR-ARCH-3's bound. Alignment runs at proxy resolution and drives the preview; the full-resolution
warp and blend run only on confirm.
**FR-MRG-8 — Model-optional.** Keypoint detection works without any model weights and better
with them, the convention `dr-segment` set. Weights that ship are recorded in `models/LICENCE.md`
before they land, under D8's compatibility test, and their absence degrades quality rather than
disabling the feature.
**FR-MRG-9 — Platform.** Desktop first. Android runs the same code within NFR-RES-2 and
FR-MRG-11, with a stated ceiling on frame count and source resolution, refused with a message,
rather than an out-of-memory kill.
**FR-MRG-10 — Where the work runs.** *(general to any merge)* Every per-pixel stage of a merge —
rendering the sources, the preview reprojection, the full-resolution warp, gain compensation, the
seam and the blend — runs on the GPU as WGSL, under ARCH §6.4. The stages that are not per-pixel
— keypoint detection at proxy resolution, descriptor matching, and the rotation solve over a few
parameters per frame — run on the CPU, and the specification says so rather than leaving it to be
"moved later".
The per-pixel stages are the whole cost, and the tablet is where the cost is paid: the output is
larger than any single photograph the pipeline has rendered, and a CPU blend of it would take
minutes there. The CPU stages are bounded by frame count, not by output size — detection is once
per frame at 1024 px, on the same runtime faces and masks use — and moving a small model's
convolutions to hand-written WGSL is real work for no visible gain. Seam finding is the one
classic stage that resists the GPU; the seam algorithm is chosen for the GPU, not for the paper.
**FR-MRG-11 — Tiled in output space.** *(general to any merge)* No stage may hold the composite as
one texture, on any platform. The warp, seam, blend and encode proceed in output-space chunks, each
pulling only the source tiles that project into it, so the working set is one chunk plus its
sources' tiles regardless of how large the composite is.
Two facts force this before memory does. `max_texture_dimension_2d` is 8192 on many mobile GPUs
and 16384 on desktop, and a three-row panorama is routinely 20 000 px wide — the composite would
not fit a texture even with the memory to spare. And five 24 MP frames at the working precision
are ~1 GB together, which the tablet does not have. ARCH §6.2 applies to the composite as it
applies to a source, and retrofitting it would be the rewrite it warns about.
**Non-goals, fixed now.** No translation solve or parallax correction — seam placement is the
tool for a hand-held set, and a photograph with real parallax is not a panorama. No HDR panorama
in one operation until HDR merge exists on its own. No live re-stitch: a different projection or
crop after the fact is a new file, not an edit. No video.
---
## 4. Non-functional requirements ## 4. Non-functional requirements
### 4.1 Performance targets ### 4.1 Performance targets
@@ -1719,6 +1827,11 @@ the class where it was expressed.
Everything the photographer produced or navigated to is a different matter, and none of it may be Everything the photographer produced or navigated to is a different matter, and none of it may be
touched by a resize. That is the list in the criterion, and it is the testable half. touched by a resize. That is the list in the criterion, and it is the testable half.
**NFR-MRG-1 — Merge latency.** For five 24 MP frames: the alignment preview (FR-MRG-7) within
5 s on the reference desktop and 15 s on the reference tablet, of which keypoint detection is at
most 1 s per frame on the tablet's CPU; the full merge written to disk within 60 s on the desktop.
The tablet's full-merge figure is set by S15, not guessed here.
**Performance regressions fail the build.** §8's benchmark suite runs per-commit; a regression **Performance regressions fail the build.** §8's benchmark suite runs per-commit; a regression
beyond a stated tolerance is a build failure, not a notification. Performance work rots otherwise. beyond a stated tolerance is a build failure, not a notification. Performance work rots otherwise.
@@ -1765,6 +1878,11 @@ means a second implementation of the operations. ARCH §6.4 stands as written, a
fallback clause has been reworded to match. A second full pipeline was the alternative, and it was fallback clause has been reworded to match. A second full pipeline was the alternative, and it was
declined for the reason ARCH §6.4 gives: the GPU path is the product, not an optimisation of it. declined for the reason ARCH §6.4 gives: the GPU path is the product, not an optimisation of it.
**NFR-MRG-2 — Reproducible merges.** The same sources, the same settings and the same device
produce a byte-identical composite. Across devices the comparison is tolerance-based, calibrated
as S9 calibrates R1: the merge is float work end to end, and ARCH §6.13's bit-identity applies to
integer state only.
### 4.3 Resource behaviour ### 4.3 Resource behaviour
**NFR-RES-1 — Bounded memory.** Memory use is bounded and configurable, independent of catalog **NFR-RES-1 — Bounded memory.** Memory use is bounded and configurable, independent of catalog
@@ -2001,6 +2119,7 @@ Rationale, evidence, and the eliminated alternatives are recorded in
| D10 | Interface strategy | One adaptive UI, tablet + desktop | | D10 | Interface strategy | One adaptive UI, tablet + desktop |
| D11 | Product positioning | Culling-first differentiator; see below | | D11 | Product positioning | Culling-first differentiator; see below |
| D12 | Scope versus pace | **DECIDED 2026-09-19** — settled by events; full scope stands, no v1 date | | D12 | Scope versus pace | **DECIDED 2026-09-19** — settled by events; full scope stands, no v1 date |
| D18 | Derived images | **DECIDED 2026-09-19** — a merge writes a new source file; no multi-source Version |
### D11 — product positioning ### D11 — product positioning
@@ -2195,6 +2314,9 @@ collides with four things this document says.
priced before starting), trashing B must know A depends on it (FR-CAT-15), and `Version::merge` priced before starting), trashing B must know A depends on it (FR-CAT-15), and `Version::merge`
has never seen a cross-reference. Reopening ARCH §6.3 for this reopens it for the three deferred has never seen a cross-reference. Reopening ARCH §6.3 for this reopens it for the three deferred
rows at once, which is the argument for doing it once and properly rather than for this alone. rows at once, which is the argument for doing it once and properly rather than for this alone.
*D18 has since done it once, the other way:* a merge writes a new file and ARCH §6.3 stands.
That leaves this case as the only one that would still need a cross-reference — the output is
frame A, not a new file — so the cost above is now this feature's alone to justify.
3. **Tone.** The source patch goes through A's chain, not B's — B demosaiced and run through A's 3. **Tone.** The source patch goes through A's chain, not B's — B demosaiced and run through A's
parameters to the head of the detail chain, then sampled. A second small pipeline at proxy parameters to the head of the detail chain, then sampled. A second small pipeline at proxy
resolution; a second full demosaic at export. The heal hides lighting drift between frames; it resolution; a second full demosaic at export. The heal hides lighting drift between frames; it
@@ -2211,6 +2333,28 @@ Deferred under D12 until FR-CULL-8a exists and spot removal's disc has, in its o
finished and used. The first cut, when it comes, is "clone from a neighbouring frame" as a spot finished and used. The first cut, when it comes, is "clone from a neighbouring frame" as a spot
source; the face-aware proposal is a layer over that. source; the face-aware proposal is a layer over that.
### D18 — derived images · **DECIDED 2026-09-19**
**A merge produces a new source file, not a multi-source Version.** The composite is written
beside its sources (FR-MRG-3), gets its own sidecar and content-hash identity, and is from then on
an ordinary `Image`: developed, synced, exported and trashed like any other. Its sidecar carries
`derived_from` (FR-MRG-6) as *provenance*, not as a *dependency* — trashing a source does not break
the composite, and rendering it needs nothing but itself.
This is the ARCH §6.3 question §7 has been keeping open for panorama, HDR and focus stacking,
answered once for all three. The alternative — a Version whose inputs are several other images,
rendered live — was priced in D17: sidecar cross-references, `Version::merge` seeing a reference
for the first time, FR-CAT-15 knowing that trashing B breaks A, FR-NC-6c requiring every source
present and priced before A can render, and an export that decodes N files. Every one of those is
a change to a subsystem that works today, and none of them buys the photographer anything a file
does not. It is also what Lightroom does, and the audience (D11) knows it.
What it forecloses, so that it reads as a decision: re-merging with different settings is a new
file, not an edit to the old one; and the composite occupies disk — a five-frame panorama is a
100–200 MB file, which the photographer chose to make. D17 narrows accordingly to the one case
where the output is still frame A, and inherits nothing from this decision but the provenance
rule.
### D16 — plugin licensing · **OPEN, post-v1** ### D16 — plugin licensing · **OPEN, post-v1**
> Deferred with §3.10 on 2026-09-19. Still to be answered before the format is published as > Deferred with §3.10 on 2026-09-19. Still to be answered before the format is published as
@@ -2248,8 +2392,9 @@ note where deferring now constrains the design later.
| Deferred | Note | | Deferred | Note |
|---|---| |---|---|
| Tethered shooting | — | | Tethered shooting | — |
| Panorama and HDR merge | **Keep the schema open** — these produce images derived from multiple sources, which ARCH §6.3's single-source `Image` cannot express. | | ~~Panorama~~ | **Undeferred 2026-09-19** — §3.11 (FR-MRG-1 … FR-MRG-11), under D18, which answers the schema question this row was holding open: a merge is a new source file, and ARCH §6.3's single-source `Image` does not change. |
| Focus stacking | Same provenance consideration. | | HDR merge | Deferred. D18 answers the data model; the merge itself — exposure alignment, ghost handling, the tone of the result — is not specified. FR-MRG-3, 5, 6, 7, 10 and 11 are written to be general to it. |
| Focus stacking | Deferred, on the same terms as HDR merge. |
| Cross-frame face repair ("best take") | A face from a neighbouring frame of the same burst, aligned by its landmarks and blended by FR-DEV-8's heal. Mechanically a spot whose source is another photograph; **the same multi-source schema question as the two rows above, arriving early** — D17. Deferred rather than refused, with its non-goals fixed now: never automatic, geometry not corrected, source frame declared in sidecar, history and export. | | Cross-frame face repair ("best take") | A face from a neighbouring frame of the same burst, aligned by its landmarks and blended by FR-DEV-8's heal. Mechanically a spot whose source is another photograph; **the same multi-source schema question as the two rows above, arriving early** — D17. Deferred rather than refused, with its non-goals fixed now: never automatic, geometry not corrected, source frame declared in sidecar, history and export. |
| Gaze and eye-contact estimation | Every open gaze model read on 2026-09-19 is trained on Gaze360, MPIIGaze or ETH-XGaze, all research-only, and Gaze360's licence restricts *models trained on it* by name — the InsightFace situation again (D13). Head pose from the five landmarks is the proxy (FR-CULL-8a). Revisit when weights with a clean data chain exist; iris offset within the eye crop is the licence-free fallback if the proxy proves too weak. | | Gaze and eye-contact estimation | Every open gaze model read on 2026-09-19 is trained on Gaze360, MPIIGaze or ETH-XGaze, all research-only, and Gaze360's licence restricts *models trained on it* by name — the InsightFace situation again (D13). Head pose from the five landmarks is the proxy (FR-CULL-8a). Revisit when weights with a clean data chain exist; iris offset within the eye crop is the licence-free fallback if the proxy proves too weak. |
| Print layout | — | | Print layout | — |
@@ -2324,6 +2469,8 @@ stacks.
| **S13** | **Slint accessibility on Android:** verify TalkBack exposure of names, roles, and values | Whether NFR-A11Y-2 is achievable in the chosen toolkit | NFR-A11Y-2 | | **S13** | **Slint accessibility on Android:** verify TalkBack exposure of names, roles, and values | Whether NFR-A11Y-2 is achievable in the chosen toolkit | NFR-A11Y-2 |
| **S14** | **Face pipeline in Rust, on a real personal library:** run a detector plus an embedder over ~2,000 images through a Rust ONNX runtime, at proxy resolution, on the reference desktop. Measure per-image latency, cluster purity against hand-labelled truth, and fit the FR-CULL-9 calibration to see whether it converges on a library-sized sample. **Resolve the model licence question before writing any of it** | Whether §3.9.1 is buildable without breaking the pure-Rust dependency policy, and whether the accuracy is worth the subsystem | D13, FR-CULL-8, FR-CULL-9 | | **S14** | **Face pipeline in Rust, on a real personal library:** run a detector plus an embedder over ~2,000 images through a Rust ONNX runtime, at proxy resolution, on the reference desktop. Measure per-image latency, cluster purity against hand-labelled truth, and fit the FR-CULL-9 calibration to see whether it converges on a library-sized sample. **Resolve the model licence question before writing any of it** | Whether §3.9.1 is buildable without breaking the pure-Rust dependency policy, and whether the accuracy is worth the subsystem | D13, FR-CULL-8, FR-CULL-9 |
| **S15** | **Panorama pre-conditions, in order of what can kill it:** (1) write a linear DNG with the `tiff` crate and read it back through rawler — decides FR-MRG-3's container; (2) export XFeat to ONNX at a fixed 1024 px input and load it under tract with zero unsupported operators — the F6 check `segmentation.md` records, and the licence read first; (3) tap the working-space texture after lens correction and before tone, and confirm it carries what FR-MRG-2 asks for; (4) a tiled multi-band blend of a 100 MP output on the reference tablet, and XFeat's per-frame time on its CPU — the two halves of NFR-MRG-1 | Whether §3.11 is buildable on the pipeline as it stands, and what the tablet figure is | D18, FR-MRG-2, FR-MRG-3, FR-MRG-8, FR-MRG-11, NFR-MRG-1 |
### Why this order ### Why this order
**S1, S2, and S10 are the three that can invalidate the architecture.** S1 and S2 test ARCH §6.1 — the **S1, S2, and S10 are the three that can invalidate the architecture.** S1 and S2 test ARCH §6.1 — the
@@ -2354,4 +2501,6 @@ Android GPU vendors.
- **Demosaic** — reconstructing full RGB from a colour-filter-array sensor capture. - **Demosaic** — reconstructing full RGB from a colour-filter-array sensor capture.
- **CFA** — colour filter array (Bayer, X-Trans). - **CFA** — colour filter array (Bayer, X-Trans).
- **Sidecar** — a small file alongside the source holding edit metadata. - **Sidecar** — a small file alongside the source holding edit metadata.
- **Composite** — an image produced by a merge (§3.11) from several sources; a source file in its own right under D18.
- **Merge** — an operation that produces a composite: panorama, HDR merge, focus stacking.
- **Pixel pipeline** — the ordered chain of processing stages from sensor data to output. - **Pixel pipeline** — the ordered chain of processing stages from sensor data to output.
+94 -81
View File
File diff suppressed because one or more lines are too long