S15.3: the camera-space tap is uniforms, not structure — and FR-MRG-2 moves below the profile

The fused chain, as operation.rs's tests fix it, is warp → as-shot white
balance → operations → base curve → camera matrix → store. LinearWorking
stores after the matrix, so the existing linear tap carries the body's
base curve, and a composite stitched from it and developed as an
unprofiled body would render that curve twice.

FR-MRG-2 therefore stitches camera-linear RGB — after the warp, before
white balance, curve and matrix — and the composite carries the first
source's body, matrices and as-shot neutral so its own develop applies
the profile once. The composer already makes this a uniform question:
white balance, matrix and the curve flag are reserved uniforms, so the
tap is a compose entry with no operations and a render entry that fills
them neutral. panorama.md §5.1 states the shape and asks for f32 buffers.
This commit is contained in:
2026-09-19 15:24:10 +02:00
parent e4b6b6c935
commit 7e6b25b21b
2 changed files with 52 additions and 12 deletions
+37 -2
View File
@@ -50,7 +50,7 @@ reused:
| Exists | Where | Reused for | | Exists | Where | Reused for |
|---|---|---| |---|---|---|
| Render a source to scene-linear on the GPU | `dr-gpu` demosaic → camera profile → working space | FR-MRG-2's input, once a tap after lens correction and before tone exists (S15.3) | | Render a source through the fused pass, with a linear f16 output mode | `dr-gpu` demosaic → `AdjustPass`, `OutputMode::LinearWorking` | FR-MRG-2's camera-space input, as a compose entry with no operations and the profile uniforms neutral (S15.3) |
| Tiled rendering with a priority scheduler | ARCH §5.3 | Pulling source tiles on demand into an output chunk (§5 below) | | Tiled rendering with a priority scheduler | ARCH §5.3 | Pulling source tiles on demand into an output chunk (§5 below) |
| A non-CFA source entering the pipeline | `Demosaicer::from_rgba8` | The composite's decode path, if the container is a TIFF (S15.1) | | A non-CFA source entering the pipeline | `Demosaicer::from_rgba8` | The composite's decode path, if the container is a TIFF (S15.1) |
| DNG matrices read through rawler | `dr-decode::profile` | The composite's decode path, if the container is a DNG | | DNG matrices read through rawler | `dr-decode::profile` | The composite's decode path, if the container is a DNG |
@@ -70,7 +70,7 @@ FR-MRG-10 states the rule; this is the table it was written from.
| Stage | Cost shape | Runs on | Why | | Stage | Cost shape | Runs on | Why |
|---|---|---|---| |---|---|---|---|
| Source to scene-linear | per pixel, full res | GPU, the existing pipeline | It *is* the pipeline | | Source to camera-linear | per pixel, full res | GPU, the existing pipeline | It *is* the pipeline, stopped early |
| Keypoint detection | once per frame, at 1024 px | CPU, tract (NEON on the tablet) | Bounded by frame count, not output size. Same runtime faces and masks use. Hand-written WGSL convolutions for a model that runs five times would be work with no visible gain. | | Keypoint detection | once per frame, at 1024 px | CPU, tract (NEON on the tablet) | Bounded by frame count, not output size. Same runtime faces and masks use. Hand-written WGSL convolutions for a model that runs five times would be work with no visible gain. |
| Descriptor matching | K² × D per pair | CPU, SIMD | 2048² × 64 × 10 pairs ≈ 3 GFLOP — tens of milliseconds | | Descriptor matching | K² × D per pair | CPU, SIMD | 2048² × 64 × 10 pairs ≈ 3 GFLOP — tens of milliseconds |
| Rotation solve, bundle adjustment | 3N + 1 parameters, Levenberg–Marquardt | CPU | Microseconds. Not parallel work. | | Rotation solve, bundle adjustment | 3N + 1 parameters, Levenberg–Marquardt | CPU | Microseconds. Not parallel work. |
@@ -125,6 +125,41 @@ neutral graph at zoom 1 and gets the same caching every other consumer does.
A tile pulled for one chunk is usually needed by the neighbouring chunk, and A tile pulled for one chunk is usually needed by the neighbouring chunk, and
stays hot for it. stays hot for it.
### 5.1 The tap — S15.3, answered by reading the composer
The fused shader's order, fixed by `operation.rs`'s own tests: warp → as-shot
white balance → operations → base curve → camera matrix → store. The store is
either the display encode or, in `OutputMode::LinearWorking`, an unclipped
`rgba16float` of linear sRGB. That mode exists for the detail stage and is
selected from the operations, never by a caller flag, so that a shader and
the texture bound to it cannot disagree.
The merge wants the values *before* the curve and matrix (FR-MRG-2), and the
composer already makes that a matter of uniforms rather than structure: the
white balance, the matrix and the curve's active flag are all in the reserved
uniform block, and a fused pass with no operations, `as_shot_wb = 1`,
`cam_to_srgb = I` and `base_curve_last.z = 0` stores exactly camera-linear
RGB after the warp. So the tap is:
- `EditGraph::compose_camera_linear()` — the `LinearWorking` tail with an
empty operation list and identity framing, paired by name with
- `AdjustPass::render_camera_linear()` — binds the f16 target, fills the
reserved uniforms neutral instead of from the source, returns the texture,
- and a float readback beside the existing 8-bit one.
Nothing in the chain moves. **Precision:** the tap and every chunk buffer
after it should be `rgba32float`, not f16. A 14-bit sensor has 16 384 steps
to white; f16 has 2 048 in the top octave, and a composite that is going to
be re-developed deserves the sensor's precision. The cost is 2× on buffers
FR-MRG-11 already bounds.
**What the DNG carries as a consequence:** the first source's `Make`,
`Model` and `UniqueCameraModel` — so `base_curve::for_body` finds the 6D's
curve — its `ColorMatrix1`/`2` with illuminants, and its `AsShotNeutral`. The
composite then develops through the same profile as its sources, applied
once. The spike's 64 × 48 file (§8) already carries the matrix and neutral;
the body name is a string.
## 6. The keypoint model ## 6. The keypoint model
FR-MRG-8: works without weights, better with them. The licence read comes FR-MRG-8: works without weights, better with them. The licence read comes
+15 -10
View File
@@ -1694,18 +1694,23 @@ which is the workflow break FR-DEV-8 was added to close for dust. It is also the
three §7 merges, and the one whose alignment problem is smallest — a rotation about one point, three §7 merges, and the one whose alignment problem is smallest — a rotation about one point,
with no depth to recover — so it is where the shared machinery is built. with no depth to recover — so it is where the shared machinery is built.
**FR-MRG-2 — What is stitched.** Each source enters the merge at develop-neutral scene-linear: **FR-MRG-2 — What is stitched.** Each source enters the merge in **camera space**: after black
after black and white levels, demosaic, camera profile and lens distortion correction, before any and white levels, demosaic and lens distortion correction, and before everything else — no white
tone or colour adjustment, with one white balance — the first frame's — applied to all. The balance, no base curve, no camera matrix, no edit. The composite carries the first source's body,
sources' own edits are not baked in. The composite is developed afterwards as if it were a new colour matrix and as-shot neutral, so that it is developed afterwards exactly as one of its
RAW. sources would be: the camera profile, the white balance and every operation in §3.3 are applied
once, to the composite, in its own develop.
This is the clause that decides what the output *is*. Stitching the rendered edits is what a JPEG This is the clause that decides what the output *is*. Stitching the rendered edits is what a JPEG
stitcher does; the result cannot be re-developed, and any difference between the frames' edits stitcher does; the result cannot be re-developed, and any difference between the frames' edits
becomes a seam. Stitching neutral pixels produces something that behaves like a photograph the becomes a seam. Stitching camera-space pixels produces a photograph the camera could have taken,
camera could have taken, and every develop operation in §3.3 then applies to it once, not five and nothing is applied twice. The cut sits *below* the profile, not above it, for a reason S15.3
times. Lens correction sits above the cut because a distorted frame does not align; white balance found in the pipeline: the base curve is part of the profile (FR-DEV-3e) and is applied to every
sits above it because the scalars must agree across frames or the overlaps do not match. frame of a known body, so a composite that baked it in and then developed as one would render the
curve twice. Lens correction alone sits above the cut, because a distorted frame does not align.
White balance sits below it because the sensor saw the same light in every frame: un-balanced
camera RGB agrees across the overlaps whether or not the camera's auto white balance drifted, and
the balanced values would not.
**FR-MRG-3 — The output file.** *(general to any merge)* Scene-linear, at least 16 bits per **FR-MRG-3 — The output file.** *(general to any merge)* Scene-linear, at least 16 bits per
channel, in a wide gamut with the colour transform resolved or the profile carried, with capture channel, in a wide gamut with the colour transform resolved or the profile carried, with capture
@@ -2469,7 +2474,7 @@ stacks.
| **S13** | **Slint accessibility on Android:** verify TalkBack exposure of names, roles, and values | Whether NFR-A11Y-2 is achievable in the chosen toolkit | NFR-A11Y-2 | | **S13** | **Slint accessibility on Android:** verify TalkBack exposure of names, roles, and values | Whether NFR-A11Y-2 is achievable in the chosen toolkit | NFR-A11Y-2 |
| **S14** | **Face pipeline in Rust, on a real personal library:** run a detector plus an embedder over ~2,000 images through a Rust ONNX runtime, at proxy resolution, on the reference desktop. Measure per-image latency, cluster purity against hand-labelled truth, and fit the FR-CULL-9 calibration to see whether it converges on a library-sized sample. **Resolve the model licence question before writing any of it** | Whether §3.9.1 is buildable without breaking the pure-Rust dependency policy, and whether the accuracy is worth the subsystem | D13, FR-CULL-8, FR-CULL-9 | | **S14** | **Face pipeline in Rust, on a real personal library:** run a detector plus an embedder over ~2,000 images through a Rust ONNX runtime, at proxy resolution, on the reference desktop. Measure per-image latency, cluster purity against hand-labelled truth, and fit the FR-CULL-9 calibration to see whether it converges on a library-sized sample. **Resolve the model licence question before writing any of it** | Whether §3.9.1 is buildable without breaking the pure-Rust dependency policy, and whether the accuracy is worth the subsystem | D13, FR-CULL-8, FR-CULL-9 |
| **S15** | **Panorama pre-conditions, in order of what can kill it:** (1) write a linear DNG with the `tiff` crate and read it back through rawler — decides FR-MRG-3's container; (2) export XFeat to ONNX at a fixed 1024 px input and load it under tract with zero unsupported operators — the F6 check `segmentation.md` records, and the licence read first; (3) tap the working-space texture after lens correction and before tone, and confirm it carries what FR-MRG-2 asks for; (4) a tiled multi-band blend of a 100 MP output on the reference tablet, and XFeat's per-frame time on its CPU — the two halves of NFR-MRG-1 | Whether §3.11 is buildable on the pipeline as it stands, and what the tablet figure is | D18, FR-MRG-2, FR-MRG-3, FR-MRG-8, FR-MRG-11, NFR-MRG-1 | | **S15** | **Panorama pre-conditions, in order of what can kill it:** (1) write a linear DNG with the `tiff` crate and read it back through rawler — decides FR-MRG-3's container; (2) export XFeat to ONNX at a fixed 1024 px input and load it under tract with zero unsupported operators — the F6 check `segmentation.md` records, and the licence read first; (3) find where in the fused chain FR-MRG-2's camera-space tap sits and what it costs to expose; (4) a tiled multi-band blend of a 100 MP output on the reference tablet, and XFeat's per-frame time on its CPU — the two halves of NFR-MRG-1 | Whether §3.11 is buildable on the pipeline as it stands, and what the tablet figure is | D18, FR-MRG-2, FR-MRG-3, FR-MRG-8, FR-MRG-11, NFR-MRG-1 |
### Why this order ### Why this order