Specify panorama merging: §3.11, D18, S15, and the design in panorama.md
A merge writes a new source file beside its sources (D18) rather than a multi-source Version, which answers the schema question §7 had been holding open for panorama, HDR merge and focus stacking together. The panorama is undeferred as FR-MRG-1 … 11; the other two stay in §7 with their data model decided. FR-MRG-10 and 11 fix where the work runs — every per-pixel stage on the GPU, the composite never held as one texture — because the output exceeds max_texture_dimension_2d before it exceeds memory. panorama.md carries the stage table, the chunked output driver, the model licences and the porting sources. S15 gates all of it. Coverage falls from 83.0% to 77.2%: thirteen requirements entered with no code, and outstanding.md §11 says so.
This commit is contained in:
+14
-1
@@ -455,7 +455,20 @@ whether something *should* be built — which is the opposite of the order §9 a
|
||||
|
||||
---
|
||||
|
||||
## 11. D12, which governed all of the above
|
||||
## 11. Merging — specified 2026-09-19, nothing built
|
||||
|
||||
§3.11 was written on 2026-09-19 under D18, undeferring the panorama from §7 and leaving HDR merge
|
||||
and focus stacking there with their data model decided. Eleven `FR-MRG` clauses and two `NFR-MRG`
|
||||
figures entered the register at once with no code behind any of them, which is why the coverage
|
||||
figure fell from 83.0% to 77.2% on the same day — a specification, not a regression.
|
||||
|
||||
[panorama.md](panorama.md) is the design, and its §10 is the order of work. Nothing starts before
|
||||
**S15**: whether rawler reads back a linear DNG the application writes, whether XFeat loads under
|
||||
tract at a fixed shape, whether the working-space texture can be tapped where FR-MRG-2 needs it,
|
||||
and what a chunked blend of a 100 MP composite costs on the tablet. The first two are a day each
|
||||
and either can change the design, which is the reason they come first.
|
||||
|
||||
## 12. D12, which governed all of the above
|
||||
|
||||
> **Decided 2026-09-19, by events.** The scope stands as calibrated, v1 has no date, and `(post-v1)`
|
||||
> in §7 is the one way a clause leaves the count — used for the plugin API and nothing else. The
|
||||
|
||||
@@ -0,0 +1,211 @@
|
||||
# Panorama
|
||||
|
||||
**Status:** Draft · 2026-09-19
|
||||
**Companion to:** [requirements.md](requirements.md) §3.11 FR-MRG-1 … 11, D18, S15 · [architecture.md](architecture.md) §5.2, §6.2
|
||||
|
||||
The first merge (§3.11): several frames, rotated about one point, become one
|
||||
photograph. This document is how that lands on the pipeline that exists now —
|
||||
which stages, where each runs, how the composite is produced in chunks when it
|
||||
is larger than any texture or any memory, what is ported from where, and what
|
||||
the keypoint model may be under D8.
|
||||
|
||||
---
|
||||
|
||||
## 1. Why it is worth the work
|
||||
|
||||
The audience shoots panoramas and leaves the application to stitch them. That
|
||||
is the same workflow break dust was (`spot-removal.md` §1): a RAW editor that
|
||||
does everything but the one thing, and the photographer's work ends up in a
|
||||
JPEG produced by a tool that never saw the RAW.
|
||||
|
||||
It is also the merge whose alignment problem is smallest. A panorama is a
|
||||
rotation — three parameters per frame plus a focal length — with no depth to
|
||||
recover. HDR merge and focus stacking share its data model (D18) and most of
|
||||
its machinery (FR-MRG-3, 5, 6, 7, 10, 11 are written to be general); building
|
||||
the panorama first builds the shared part on the easiest geometry.
|
||||
|
||||
## 2. Non-goals
|
||||
|
||||
- **Not structure-from-motion.** No translation is solved for. A hand-held set
|
||||
with parallax gets its ghosts hidden by seam placement, and a set with real
|
||||
parallax is not a panorama. COLMAP's front end is the right mental model;
|
||||
its back end is the wrong problem.
|
||||
- **Not a multi-source Version.** D18. The composite is a file, and nothing in
|
||||
the catalog, the sidecar format or sync learns about cross-references.
|
||||
- **Not boundary fill.** Painting pixels that were never captured is the pixel
|
||||
editing §1.3 excludes. Auto-crop is the tool.
|
||||
- **Not automatic.** The tool proposes an alignment and writes nothing until
|
||||
the photographer confirms. Same rule as spot removal and D17, for the same
|
||||
reason: a merge that silently omits or misplaces a frame is the failure this
|
||||
application must not have.
|
||||
- **Not HDR-panorama in one pass.** Until HDR merge exists on its own, a
|
||||
bracketed panorama is bracketed frames merged first, then stitched.
|
||||
|
||||
## 3. What is new, precisely
|
||||
|
||||
Nearly all of it, unlike spot removal. The pipeline renders one source to one
|
||||
texture; nothing in the tree detects keypoints, estimates a rotation, warps
|
||||
into a projection, finds a seam, or blends a pyramid. What exists and is
|
||||
reused:
|
||||
|
||||
| Exists | Where | Reused for |
|
||||
|---|---|---|
|
||||
| Render a source to scene-linear on the GPU | `dr-gpu` demosaic → camera profile → working space | FR-MRG-2's input, once a tap after lens correction and before tone exists (S15.3) |
|
||||
| Tiled rendering with a priority scheduler | ARCH §5.3 | Pulling source tiles on demand into an output chunk (§5 below) |
|
||||
| A non-CFA source entering the pipeline | `Demosaicer::from_rgba8` | The composite's decode path, if the container is a TIFF (S15.1) |
|
||||
| DNG matrices read through rawler | `dr-decode::profile` | The composite's decode path, if the container is a DNG |
|
||||
| Static-shape ONNX under tract, heads decoded in Rust | `dr-segment` | The keypoint model (§6) |
|
||||
| A batch worker with its own `GpuContext`, activity row, cancel | `dr-ui::export` | FR-MRG-7 verbatim |
|
||||
| The 16-bit TIFF encoder with metadata sub-IFDs | `dr-export::encode` | FR-MRG-3's writer, extended to linear samples |
|
||||
| Multi-select in the grid | `collections_ui::selected` | The entry point |
|
||||
|
||||
New: a `core/dr-pano` crate holding the geometry (keypoints, matching, the
|
||||
rotation solve), a set of WGSL passes in `dr-gpu` (reprojection, gain,
|
||||
seam, pyramid blend), the chunked output driver, the container writer, and
|
||||
the dialog.
|
||||
|
||||
## 4. The stages, and where each runs
|
||||
|
||||
FR-MRG-10 states the rule; this is the table it was written from.
|
||||
|
||||
| Stage | Cost shape | Runs on | Why |
|
||||
|---|---|---|---|
|
||||
| Source to scene-linear | per pixel, full res | GPU, the existing pipeline | It *is* the pipeline |
|
||||
| Keypoint detection | once per frame, at 1024 px | CPU, tract (NEON on the tablet) | Bounded by frame count, not output size. Same runtime faces and masks use. Hand-written WGSL convolutions for a model that runs five times would be work with no visible gain. |
|
||||
| Descriptor matching | K² × D per pair | CPU, SIMD | 2048² × 64 × 10 pairs ≈ 3 GFLOP — tens of milliseconds |
|
||||
| Rotation solve, bundle adjustment | 3N + 1 parameters, Levenberg–Marquardt | CPU | Microseconds. Not parallel work. |
|
||||
| Preview reprojection | per pixel, proxy res | GPU, interactive | Projection and horizon changes re-warp N proxies at frame rate |
|
||||
| Full-resolution warp | per output pixel | GPU, chunked (§5) | The heaviest thing in the application |
|
||||
| Gain compensation | per overlap region | GPU reduction, then N scalars | Sums, on the histogram pass's pattern (ARCH §5.5) |
|
||||
| Seam finding | per overlap pixel | GPU-friendly variant | Graph cut resists the GPU; a distance-transform or per-column DP seam does not. The algorithm is chosen for the GPU, not for the paper. |
|
||||
| Multi-band blend | per pixel × levels | GPU, chunked | Laplacian pyramids are separable convolutions — the detail stage's shape |
|
||||
| Encode | per pixel, once | CPU, streamed per chunk row | As export does |
|
||||
|
||||
## 5. Chunked in output space
|
||||
|
||||
FR-MRG-11 forbids holding the composite as one texture, and two facts force it
|
||||
before memory does:
|
||||
|
||||
- `max_texture_dimension_2d` is 8192 on many mobile GPUs and 16384 on desktop.
|
||||
A three-row panorama is routinely 20 000 px wide.
|
||||
- Five 24 MP frames at working precision are ~1 GB together. The tablet does
|
||||
not have it.
|
||||
|
||||
**The geometry is known before any full-resolution pixel exists.** Alignment
|
||||
runs on proxies; what comes out is a rotation per frame, a focal length, a
|
||||
projection and an output rectangle. From those, every output pixel's source
|
||||
coordinates in every frame are a closed-form function. That is what makes
|
||||
chunking simple rather than clever:
|
||||
|
||||
```
|
||||
for each output chunk C (e.g. 2048 × 2048, in output space):
|
||||
frames_in(C) = frames whose projected footprint intersects C
|
||||
for each frame F in frames_in(C):
|
||||
source tiles T(F, C) = tiles of F that project into C, plus a margin
|
||||
render T(F, C) to scene-linear through the pipeline's tile cache
|
||||
warp T(F, C) into C's coordinate frame ← GPU
|
||||
gain-correct, seam, blend within C ← GPU, with overlap
|
||||
read C back, encode its rows ← CPU, streamed
|
||||
```
|
||||
|
||||
The working set is one chunk, its per-frame warped copies, and the source
|
||||
tiles that fed them. It does not grow with the composite.
|
||||
|
||||
**The blend needs a margin.** A Laplacian pyramid of L levels reads
|
||||
2^L pixels beyond the chunk edge; a chunk is therefore rendered with a margin
|
||||
of that width and the margin discarded after the blend. Seams cross chunk
|
||||
boundaries and must agree on both sides: the seam is found once at a reduced
|
||||
resolution over the whole overlap (which fits — it is a mask, not an image),
|
||||
then upsampled into each chunk. The same is true of gain: the scalars are
|
||||
solved once from proxy-resolution overlaps and applied everywhere.
|
||||
|
||||
**Source tiles are the pipeline's tiles.** ARCH §5.3's cache keys by
|
||||
`(VersionId, tile, zoom, graph_hash_prefix)`; the merge asks for tiles of a
|
||||
neutral graph at zoom 1 and gets the same caching every other consumer does.
|
||||
A tile pulled for one chunk is usually needed by the neighbouring chunk, and
|
||||
stays hot for it.
|
||||
|
||||
## 6. The keypoint model
|
||||
|
||||
FR-MRG-8: works without weights, better with them. The licence read comes
|
||||
first (D13's lesson, S15.2).
|
||||
|
||||
| Model | Licence | Fits tract? | Position |
|
||||
|---|---|---|---|
|
||||
| **XFeat** (CVPR 2024) | Apache-2.0 | Plain convolutions, fully convolutional, the repo ships an ONNX export | **Chosen.** Fixed 1024 px input, dense heatmap and descriptor map out, NMS and top-K in Rust — the yolo26 pattern |
|
||||
| DISK | Apache-2.0 | U-Net, static | Second choice; stronger descriptors, ~3–4× the compute |
|
||||
| ALIKE | BSD-3 | Plain convolutions | Fallback if XFeat's export fails F6 |
|
||||
| ALIKED | BSD-3 | Deformable convolution in the descriptor head | Unlikely to load |
|
||||
| SuperPoint, SuperGlue, R2D2, SiLK, MASt3R | non-commercial | — | Out on licence |
|
||||
| LightGlue | Apache-2.0 | Transformer over a variable keypoint count | Not until mutual-nearest-neighbour matching fails on a real set |
|
||||
|
||||
Without weights: AKAZE (BSD, `akaze` from rust-cv), which is adequate on
|
||||
well-textured overlaps and worse on sky, repeated structure and exposure
|
||||
drift — which is where a learned detector earns its place.
|
||||
|
||||
Matching is mutual nearest neighbour with a ratio test, then RANSAC on a
|
||||
rotation model. For a panorama — one lens, near-pure rotation, 20–40 %
|
||||
overlap — that is what Hugin and OpenCV's stitcher use, and it is enough.
|
||||
|
||||
## 7. What is ported from where
|
||||
|
||||
Nothing is linked; everything is read.
|
||||
|
||||
| Source | Licence | Taken |
|
||||
|---|---|---|
|
||||
| OpenCV `modules/stitching` | Apache-2.0 | The stage layout — Brown & Lowe (2007) as a set of small classes with one job each — and the warpers' projection maths |
|
||||
| OpenPano (ppwwyyxx) | MIT (verify on read) | The estimation and bundle-adjustment maths, function by function, with outputs diffed against it |
|
||||
| enblend-enfuse | GPLv2+ | Seam-line optimisation and Burt–Adelson multi-band blending |
|
||||
| Hugin `nona` | GPLv2+ | The GLSL remapper, as the reference for the WGSL warp |
|
||||
|
||||
The golden set (§8 of the requirements) is OpenCV's stitcher on the same
|
||||
inputs: a reference output to compare against, within a tolerance calibrated
|
||||
the way S9 calibrates R1.
|
||||
|
||||
## 8. The output file
|
||||
|
||||
FR-MRG-3. Scene-linear, ≥ 16 bits, wide gamut, the first source's capture
|
||||
metadata, named from the first source with a `-pano` suffix, beside it.
|
||||
|
||||
Two containers are candidates and S15.1 decides:
|
||||
|
||||
- **Linear DNG.** `PhotometricInterpretation = LinearRaw`, three samples per
|
||||
pixel, `ColorMatrix1` carried from the first source. Re-enters through
|
||||
rawler as `Format::Dng` with no new decode path, *if* rawler reads it back.
|
||||
What Lightroom writes.
|
||||
- **Float TIFF.** `SampleFormat = IEEEFP`, 16 or 32 bits, an ICC profile for
|
||||
the working space. Needs `Format::Tiff` and a decode path, but the writer is
|
||||
the existing encoder with a different sample type, and nothing about it is
|
||||
uncertain.
|
||||
|
||||
Either way the composite enters the pipeline as a non-CFA, *linear* source —
|
||||
`from_rgba8`'s sibling with `non_linear = false` and the base curve resolved
|
||||
from the carried matrix — and is developed as any RAW is.
|
||||
|
||||
## 9. Interaction
|
||||
|
||||
- The entry is the grid's selection: two or more images, one action, "Merge
|
||||
to panorama". One image, or images from different roots, and the action
|
||||
says why it is unavailable.
|
||||
- The dialog shows the aligned proxies in the chosen projection, with the
|
||||
projection, horizon and crop controls of FR-MRG-4, and the per-frame
|
||||
residuals. A frame that failed to align is named there (FR-MRG-5), and the
|
||||
merge cannot be confirmed with it in the set.
|
||||
- Confirm starts the FR-MRG-7 job. The composite appears in the grid when the
|
||||
file is written and catalogued, beside its sources, with the merge as the
|
||||
first entry in its history.
|
||||
|
||||
## 10. Order of work
|
||||
|
||||
1. **S15**, all four, before anything else. (1) and (2) are a day each and
|
||||
either can change the design.
|
||||
2. `dr-pano`: keypoints (AKAZE first, XFeat when S15.2 passes), matching,
|
||||
RANSAC, rotation solve. Unit-tested against synthetic rotations of one
|
||||
frame, where the answer is known exactly.
|
||||
3. The working-space tap, and the preview reprojection pass. At this point the
|
||||
dialog can show an alignment.
|
||||
4. The chunked driver with a feathered blend — the whole path end to end,
|
||||
writing a file, before the blend is good.
|
||||
5. Gain, seams, multi-band.
|
||||
6. The container, the catalog entry, provenance, the history entry.
|
||||
7. Tablet: NFR-MRG-1's figure, and FR-MRG-9's ceiling.
|
||||
+151
-2
@@ -1678,6 +1678,114 @@ application starts, opens a photograph, names the responsible plugin, and contin
|
||||
|
||||
---
|
||||
|
||||
### 3.11 Merging images
|
||||
|
||||
Several photographs become one. Panorama is the first merge and the only one specified; HDR merge
|
||||
and focus stacking share its data model (D18) and are still deferred in §7. The clauses below are
|
||||
written for the panorama and, where a clause is general to any merge, say so.
|
||||
|
||||
**FR-MRG-1 — Panorama from a selection.** Two or more selected images are aligned and blended
|
||||
into one composite, which is written as a new source file per D18. The tool is never automatic:
|
||||
it proposes an alignment, the photographer sees it and confirms, and nothing is written before
|
||||
that press.
|
||||
|
||||
The stated audience (D11) shoots panoramas and currently leaves the application to stitch them,
|
||||
which is the workflow break FR-DEV-8 was added to close for dust. It is also the first of the
|
||||
three §7 merges, and the one whose alignment problem is smallest — a rotation about one point,
|
||||
with no depth to recover — so it is where the shared machinery is built.
|
||||
|
||||
**FR-MRG-2 — What is stitched.** Each source enters the merge at develop-neutral scene-linear:
|
||||
after black and white levels, demosaic, camera profile and lens distortion correction, before any
|
||||
tone or colour adjustment, with one white balance — the first frame's — applied to all. The
|
||||
sources' own edits are not baked in. The composite is developed afterwards as if it were a new
|
||||
RAW.
|
||||
|
||||
This is the clause that decides what the output *is*. Stitching the rendered edits is what a JPEG
|
||||
stitcher does; the result cannot be re-developed, and any difference between the frames' edits
|
||||
becomes a seam. Stitching neutral pixels produces something that behaves like a photograph the
|
||||
camera could have taken, and every develop operation in §3.3 then applies to it once, not five
|
||||
times. Lens correction sits above the cut because a distorted frame does not align; white balance
|
||||
sits above it because the scalars must agree across frames or the overlaps do not match.
|
||||
|
||||
**FR-MRG-3 — The output file.** *(general to any merge)* Scene-linear, at least 16 bits per
|
||||
channel, in a wide gamut with the colour transform resolved or the profile carried, with capture
|
||||
metadata from the first source. Named from the first source with a stated suffix and placed
|
||||
beside it. Where the sources' folder is not writable — a remote-only tier, a read-only mount —
|
||||
it goes where an export goes (FR-EXP-6) and the interface says so before the merge starts.
|
||||
|
||||
The container is a spike result (S15), not a requirement: a linear DNG if rawler reads back what
|
||||
the application writes, else a float TIFF with a decode path of its own. The choice is invisible
|
||||
to everything above the decoder.
|
||||
|
||||
**FR-MRG-4 — Projection and framing.** Cylindrical, spherical or perspective, chosen from the
|
||||
field of view and overridable; the horizon levelled from the estimated rotations, overridable by a
|
||||
drag; auto-crop to the largest inscribed rectangle, overridable. No boundary fill: painting pixels
|
||||
that were never captured is the pixel editing §1.3 excludes.
|
||||
|
||||
**FR-MRG-5 — Honesty of failure.** *(general to any merge)* A frame that cannot be aligned is
|
||||
named, with why — too few matches, no overlap with any other frame, a residual above the stated
|
||||
bound — and the merge stops. Never a silent drop, never a best-effort composite with a frame
|
||||
missing.
|
||||
|
||||
The same rule as `spot-removal.md`'s and D17's: a tool that quietly alters or omits part of a
|
||||
photograph is the failure this application must not have, and here the omission would be an
|
||||
entire frame.
|
||||
|
||||
**FR-MRG-6 — Provenance.** *(general to any merge)* The composite's sidecar carries
|
||||
`derived_from`: the content hashes of its sources in order, and the merge parameters. The history
|
||||
records the merge as the first entry, and export metadata declares the composite as one. Sources
|
||||
trashed later leave the list dangling; the panel says so and nothing is blocked.
|
||||
|
||||
Provenance, not dependency. The composite renders from itself alone; `derived_from` exists so the
|
||||
photographer, and anyone they hand the file to, can see what it is made of. The audience is
|
||||
RAW-literate and a composite declares itself (D17). Nothing here specifies C2PA.
|
||||
|
||||
**FR-MRG-7 — Execution.** *(general to any merge)* A background job on the pattern FR-EXP-7
|
||||
established: its own thread, its own `GpuContext`, a row in the activity panel, cancellable with
|
||||
NFR-ARCH-3's bound. Alignment runs at proxy resolution and drives the preview; the full-resolution
|
||||
warp and blend run only on confirm.
|
||||
|
||||
**FR-MRG-8 — Model-optional.** Keypoint detection works without any model weights and better
|
||||
with them, the convention `dr-segment` set. Weights that ship are recorded in `models/LICENCE.md`
|
||||
before they land, under D8's compatibility test, and their absence degrades quality rather than
|
||||
disabling the feature.
|
||||
|
||||
**FR-MRG-9 — Platform.** Desktop first. Android runs the same code within NFR-RES-2 and
|
||||
FR-MRG-11, with a stated ceiling on frame count and source resolution, refused with a message,
|
||||
rather than an out-of-memory kill.
|
||||
|
||||
**FR-MRG-10 — Where the work runs.** *(general to any merge)* Every per-pixel stage of a merge —
|
||||
rendering the sources, the preview reprojection, the full-resolution warp, gain compensation, the
|
||||
seam and the blend — runs on the GPU as WGSL, under ARCH §6.4. The stages that are not per-pixel
|
||||
— keypoint detection at proxy resolution, descriptor matching, and the rotation solve over a few
|
||||
parameters per frame — run on the CPU, and the specification says so rather than leaving it to be
|
||||
"moved later".
|
||||
|
||||
The per-pixel stages are the whole cost, and the tablet is where the cost is paid: the output is
|
||||
larger than any single photograph the pipeline has rendered, and a CPU blend of it would take
|
||||
minutes there. The CPU stages are bounded by frame count, not by output size — detection is once
|
||||
per frame at 1024 px, on the same runtime faces and masks use — and moving a small model's
|
||||
convolutions to hand-written WGSL is real work for no visible gain. Seam finding is the one
|
||||
classic stage that resists the GPU; the seam algorithm is chosen for the GPU, not for the paper.
|
||||
|
||||
**FR-MRG-11 — Tiled in output space.** *(general to any merge)* No stage may hold the composite as
|
||||
one texture, on any platform. The warp, seam, blend and encode proceed in output-space chunks, each
|
||||
pulling only the source tiles that project into it, so the working set is one chunk plus its
|
||||
sources' tiles regardless of how large the composite is.
|
||||
|
||||
Two facts force this before memory does. `max_texture_dimension_2d` is 8192 on many mobile GPUs
|
||||
and 16384 on desktop, and a three-row panorama is routinely 20 000 px wide — the composite would
|
||||
not fit a texture even with the memory to spare. And five 24 MP frames at the working precision
|
||||
are ~1 GB together, which the tablet does not have. ARCH §6.2 applies to the composite as it
|
||||
applies to a source, and retrofitting it would be the rewrite it warns about.
|
||||
|
||||
**Non-goals, fixed now.** No translation solve or parallax correction — seam placement is the
|
||||
tool for a hand-held set, and a photograph with real parallax is not a panorama. No HDR panorama
|
||||
in one operation until HDR merge exists on its own. No live re-stitch: a different projection or
|
||||
crop after the fact is a new file, not an edit. No video.
|
||||
|
||||
---
|
||||
|
||||
## 4. Non-functional requirements
|
||||
|
||||
### 4.1 Performance targets
|
||||
@@ -1719,6 +1827,11 @@ the class where it was expressed.
|
||||
Everything the photographer produced or navigated to is a different matter, and none of it may be
|
||||
touched by a resize. That is the list in the criterion, and it is the testable half.
|
||||
|
||||
**NFR-MRG-1 — Merge latency.** For five 24 MP frames: the alignment preview (FR-MRG-7) within
|
||||
5 s on the reference desktop and 15 s on the reference tablet, of which keypoint detection is at
|
||||
most 1 s per frame on the tablet's CPU; the full merge written to disk within 60 s on the desktop.
|
||||
The tablet's full-merge figure is set by S15, not guessed here.
|
||||
|
||||
**Performance regressions fail the build.** §8's benchmark suite runs per-commit; a regression
|
||||
beyond a stated tolerance is a build failure, not a notification. Performance work rots otherwise.
|
||||
|
||||
@@ -1765,6 +1878,11 @@ means a second implementation of the operations. ARCH §6.4 stands as written, a
|
||||
fallback clause has been reworded to match. A second full pipeline was the alternative, and it was
|
||||
declined for the reason ARCH §6.4 gives: the GPU path is the product, not an optimisation of it.
|
||||
|
||||
**NFR-MRG-2 — Reproducible merges.** The same sources, the same settings and the same device
|
||||
produce a byte-identical composite. Across devices the comparison is tolerance-based, calibrated
|
||||
as S9 calibrates R1: the merge is float work end to end, and ARCH §6.13's bit-identity applies to
|
||||
integer state only.
|
||||
|
||||
### 4.3 Resource behaviour
|
||||
|
||||
**NFR-RES-1 — Bounded memory.** Memory use is bounded and configurable, independent of catalog
|
||||
@@ -2001,6 +2119,7 @@ Rationale, evidence, and the eliminated alternatives are recorded in
|
||||
| D10 | Interface strategy | One adaptive UI, tablet + desktop |
|
||||
| D11 | Product positioning | Culling-first differentiator; see below |
|
||||
| D12 | Scope versus pace | **DECIDED 2026-09-19** — settled by events; full scope stands, no v1 date |
|
||||
| D18 | Derived images | **DECIDED 2026-09-19** — a merge writes a new source file; no multi-source Version |
|
||||
|
||||
### D11 — product positioning
|
||||
|
||||
@@ -2195,6 +2314,9 @@ collides with four things this document says.
|
||||
priced before starting), trashing B must know A depends on it (FR-CAT-15), and `Version::merge`
|
||||
has never seen a cross-reference. Reopening ARCH §6.3 for this reopens it for the three deferred
|
||||
rows at once, which is the argument for doing it once and properly rather than for this alone.
|
||||
*D18 has since done it once, the other way:* a merge writes a new file and ARCH §6.3 stands.
|
||||
That leaves this case as the only one that would still need a cross-reference — the output is
|
||||
frame A, not a new file — so the cost above is now this feature's alone to justify.
|
||||
3. **Tone.** The source patch goes through A's chain, not B's — B demosaiced and run through A's
|
||||
parameters to the head of the detail chain, then sampled. A second small pipeline at proxy
|
||||
resolution; a second full demosaic at export. The heal hides lighting drift between frames; it
|
||||
@@ -2211,6 +2333,28 @@ Deferred under D12 until FR-CULL-8a exists and spot removal's disc has, in its o
|
||||
finished and used. The first cut, when it comes, is "clone from a neighbouring frame" as a spot
|
||||
source; the face-aware proposal is a layer over that.
|
||||
|
||||
### D18 — derived images · **DECIDED 2026-09-19**
|
||||
|
||||
**A merge produces a new source file, not a multi-source Version.** The composite is written
|
||||
beside its sources (FR-MRG-3), gets its own sidecar and content-hash identity, and is from then on
|
||||
an ordinary `Image`: developed, synced, exported and trashed like any other. Its sidecar carries
|
||||
`derived_from` (FR-MRG-6) as *provenance*, not as a *dependency* — trashing a source does not break
|
||||
the composite, and rendering it needs nothing but itself.
|
||||
|
||||
This is the ARCH §6.3 question §7 has been keeping open for panorama, HDR and focus stacking,
|
||||
answered once for all three. The alternative — a Version whose inputs are several other images,
|
||||
rendered live — was priced in D17: sidecar cross-references, `Version::merge` seeing a reference
|
||||
for the first time, FR-CAT-15 knowing that trashing B breaks A, FR-NC-6c requiring every source
|
||||
present and priced before A can render, and an export that decodes N files. Every one of those is
|
||||
a change to a subsystem that works today, and none of them buys the photographer anything a file
|
||||
does not. It is also what Lightroom does, and the audience (D11) knows it.
|
||||
|
||||
What it forecloses, so that it reads as a decision: re-merging with different settings is a new
|
||||
file, not an edit to the old one; and the composite occupies disk — a five-frame panorama is a
|
||||
100–200 MB file, which the photographer chose to make. D17 narrows accordingly to the one case
|
||||
where the output is still frame A, and inherits nothing from this decision but the provenance
|
||||
rule.
|
||||
|
||||
### D16 — plugin licensing · **OPEN, post-v1**
|
||||
|
||||
> Deferred with §3.10 on 2026-09-19. Still to be answered before the format is published as
|
||||
@@ -2248,8 +2392,9 @@ note where deferring now constrains the design later.
|
||||
| Deferred | Note |
|
||||
|---|---|
|
||||
| Tethered shooting | — |
|
||||
| Panorama and HDR merge | **Keep the schema open** — these produce images derived from multiple sources, which ARCH §6.3's single-source `Image` cannot express. |
|
||||
| Focus stacking | Same provenance consideration. |
|
||||
| ~~Panorama~~ | **Undeferred 2026-09-19** — §3.11 (FR-MRG-1 … FR-MRG-11), under D18, which answers the schema question this row was holding open: a merge is a new source file, and ARCH §6.3's single-source `Image` does not change. |
|
||||
| HDR merge | Deferred. D18 answers the data model; the merge itself — exposure alignment, ghost handling, the tone of the result — is not specified. FR-MRG-3, 5, 6, 7, 10 and 11 are written to be general to it. |
|
||||
| Focus stacking | Deferred, on the same terms as HDR merge. |
|
||||
| Cross-frame face repair ("best take") | A face from a neighbouring frame of the same burst, aligned by its landmarks and blended by FR-DEV-8's heal. Mechanically a spot whose source is another photograph; **the same multi-source schema question as the two rows above, arriving early** — D17. Deferred rather than refused, with its non-goals fixed now: never automatic, geometry not corrected, source frame declared in sidecar, history and export. |
|
||||
| Gaze and eye-contact estimation | Every open gaze model read on 2026-09-19 is trained on Gaze360, MPIIGaze or ETH-XGaze, all research-only, and Gaze360's licence restricts *models trained on it* by name — the InsightFace situation again (D13). Head pose from the five landmarks is the proxy (FR-CULL-8a). Revisit when weights with a clean data chain exist; iris offset within the eye crop is the licence-free fallback if the proxy proves too weak. |
|
||||
| Print layout | — |
|
||||
@@ -2324,6 +2469,8 @@ stacks.
|
||||
| **S13** | **Slint accessibility on Android:** verify TalkBack exposure of names, roles, and values | Whether NFR-A11Y-2 is achievable in the chosen toolkit | NFR-A11Y-2 |
|
||||
| **S14** | **Face pipeline in Rust, on a real personal library:** run a detector plus an embedder over ~2,000 images through a Rust ONNX runtime, at proxy resolution, on the reference desktop. Measure per-image latency, cluster purity against hand-labelled truth, and fit the FR-CULL-9 calibration to see whether it converges on a library-sized sample. **Resolve the model licence question before writing any of it** | Whether §3.9.1 is buildable without breaking the pure-Rust dependency policy, and whether the accuracy is worth the subsystem | D13, FR-CULL-8, FR-CULL-9 |
|
||||
|
||||
| **S15** | **Panorama pre-conditions, in order of what can kill it:** (1) write a linear DNG with the `tiff` crate and read it back through rawler — decides FR-MRG-3's container; (2) export XFeat to ONNX at a fixed 1024 px input and load it under tract with zero unsupported operators — the F6 check `segmentation.md` records, and the licence read first; (3) tap the working-space texture after lens correction and before tone, and confirm it carries what FR-MRG-2 asks for; (4) a tiled multi-band blend of a 100 MP output on the reference tablet, and XFeat's per-frame time on its CPU — the two halves of NFR-MRG-1 | Whether §3.11 is buildable on the pipeline as it stands, and what the tablet figure is | D18, FR-MRG-2, FR-MRG-3, FR-MRG-8, FR-MRG-11, NFR-MRG-1 |
|
||||
|
||||
### Why this order
|
||||
|
||||
**S1, S2, and S10 are the three that can invalidate the architecture.** S1 and S2 test ARCH §6.1 — the
|
||||
@@ -2354,4 +2501,6 @@ Android GPU vendors.
|
||||
- **Demosaic** — reconstructing full RGB from a colour-filter-array sensor capture.
|
||||
- **CFA** — colour filter array (Bayer, X-Trans).
|
||||
- **Sidecar** — a small file alongside the source holding edit metadata.
|
||||
- **Composite** — an image produced by a merge (§3.11) from several sources; a source file in its own right under D18.
|
||||
- **Merge** — an operation that produces a composite: panorama, HDR merge, focus stacking.
|
||||
- **Pixel pipeline** — the ordered chain of processing stages from sensor data to output.
|
||||
|
||||
+94
-81
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user