docs/ had 26 developer documents flat beside the manual, and the two audiences are very differently sized: most readers want the manual and the gesture reference, a few want the register, the designs and the measurements. The manual and gestures.md stay at the top; everything for someone changing the code moves to docs/dev/, and the two documents that name their own successors — the v0.1 milestone and the UI-refinement plan — go to docs/dev/archive/ rather than being deleted, since both are still cited. docs/README.md is the index, users first. Every reference follows: code comments, Cargo manifests, the workflows, the pre-commit hook, the bench and traceability tools (which locate the repo root by docs/dev/requirements.md now), packaging, the Docker READMEs, CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level deeper and is regenerated. Links out of the moved documents into the tree gain a level; a link checker over every Markdown file finds none broken.
504 lines
29 KiB
Markdown
504 lines
29 KiB
Markdown
# Region segmentation for local masking
|
||
|
||
Spec for **S15**, the spike that decides how DarkRoom finds the boundaries a local mask snaps to.
|
||
|
||
Local adjustments (FR-DEV-3, "linear gradient, radial gradient, and brush masks") need more than
|
||
placement handles to be competitive. The interactions that matter — click to select a region, drag a
|
||
contour that clings to an edge, paint without crossing a boundary — all need the same thing
|
||
underneath: **a map of where the image's regions are.**
|
||
|
||
There are two credible ways to produce that map and they are not obviously ordered. This document
|
||
specifies both, specifies the third option of combining them, and fixes the measurements that decide
|
||
between them *before* any of them is built.
|
||
|
||
---
|
||
|
||
## 1. Why this is a spike and not a build
|
||
|
||
Three properties make the choice expensive to get wrong.
|
||
|
||
**It sets the mask representation.** If regions exist, a mask is a *set of region ids* — integers,
|
||
diffable, mergeable at node level under FR-NC-9, cheap in a sidecar. If they don't, a mask is a
|
||
raster, and rasters are none of those things. This is the decision that is expensive to retrofit;
|
||
everything else in local masking sits on top of it.
|
||
|
||
**One arm collides with a settled policy.** D13 records that every dependency choice in this project
|
||
has gone the same way — rustls over aws-lc-rs, bundled SQLite, a Rust Lensfun port, zune-jpeg — to
|
||
avoid a C dependency under the Android NDK, and names `ort` as the largest exception that policy
|
||
would tolerate. Arm B needs exactly that exception. Arm A needs no new dependency at all. That
|
||
asymmetry is not a tiebreak, it is most of the cost difference, and it should be priced honestly
|
||
rather than discovered at packaging time.
|
||
|
||
**Model licensing is a distribution blocker.** D13 already establishes this for the face pipeline,
|
||
and the same reading applies here — see §7. It is a licence-reading exercise, not a research
|
||
question, and it comes first.
|
||
|
||
---
|
||
|
||
## 2. The common interface
|
||
|
||
Both arms produce the same thing. This is what makes them comparable, and what lets the choice be
|
||
deferred behind a seam rather than baked into every consumer.
|
||
|
||
```rust
|
||
/// A partition of the image into labelled regions.
|
||
pub struct RegionField {
|
||
/// Per-pixel region id at proxy resolution. R32Uint on the GPU.
|
||
labels: Texture,
|
||
/// Per-region summary: pixel count, bounding box, mean colour,
|
||
/// and (arm B only) a semantic class id.
|
||
regions: Vec<Region>,
|
||
/// Boundary strength per adjacent region pair — the edge weight
|
||
/// the merge tree is built from and the cost field reads.
|
||
adjacency: Vec<(RegionId, RegionId, f32)>,
|
||
}
|
||
```
|
||
|
||
Two consumers sit on it, and neither knows which arm produced it:
|
||
|
||
**A cost field, for contour snapping.** Live-wire — Dijkstra from the last anchor to the cursor over
|
||
a per-pixel cost that is *low* on boundaries. The cost is a **sum of terms**, which is the property
|
||
that matters: image gradient is always available, region boundary strength is added when a
|
||
`RegionField` exists, and a semantic boundary term is added when a model is present. Each source
|
||
improves the snap without changing the interface, so the arms are not exclusive here even in
|
||
principle.
|
||
|
||
**A region set, for click selection.** Click reads the label under the cursor; the mask is
|
||
`label(px) ∈ selected`. Add and subtract are set operations on ids. No flood fill, no readback, no
|
||
iteration — the whole reason the precomputed map is worth having.
|
||
|
||
Both consumers are built once, in the spike, and shared by both arms. A comparison in which each arm
|
||
gets its own consumer measures the consumers.
|
||
|
||
---
|
||
|
||
## 3. Arm A — multiscale watershed
|
||
|
||
No model, no new dependency, deterministic, works on any image.
|
||
|
||
**Gradient.** Sobel magnitude over a perceptual luma plus chroma distance, not camera-space RGB —
|
||
channel-weighted RGB gradient reads a saturated red edge as weaker than it looks. Computed after
|
||
demosaic and denoise, before the edit graph, so an exposure change does not invalidate it.
|
||
|
||
**Pre-smoothing is not optional.** Raw watershed on a noisy file makes every grain its own basin.
|
||
A guided or bilateral pre-filter, with strength tied to the file's ISO, is part of the arm rather
|
||
than a refinement of it.
|
||
|
||
**Basins.** Each pixel points downhill to its steepest neighbour; pointer-jumping resolves every
|
||
pixel to its basin root in log passes. Two compute shaders and a dispatch loop.
|
||
|
||
**The hierarchy is the cheap part.** Build the region adjacency graph, sort edges by boundary
|
||
strength, union-find over them, and *record the merge order*. That recording is the merge tree — a
|
||
click selects a leaf, and a scroll walks up through progressively coarser merges. Textbook Kruskal on
|
||
a graph of a few thousand nodes.
|
||
|
||
That node count is why the tree build is a legitimate CPU operation: it runs on the adjacency graph,
|
||
not on pixels. Pixels stay on the GPU, the graph is CPU-side — the same split ARCH §3.4 and §6.1
|
||
already draw for the edit graph, so no exception to the no-readback rule is needed.
|
||
|
||
**Known weaknesses, to be measured rather than argued about:** over-segmentation on noise and
|
||
texture, weak boundaries where contrast is low but semantics are obvious (a pale sky meeting a pale
|
||
wall), and a granularity ladder that is geometric rather than meaningful — level 7 is *a* coarser
|
||
partition, not necessarily *the* object.
|
||
|
||
---
|
||
|
||
## 4. Arm B — semantic segmentation
|
||
|
||
YOLO26-seg pretrained on ADE20K, run through `ort`.
|
||
|
||
ADE20K's 150 classes include stuff — sky, vegetation, water, wall, road — which is a far better
|
||
vocabulary for photography than COCO's 80 thing-classes. "That patch of sky" is a class here. The
|
||
nano variant is ~1.6M parameters, which is genuinely mobile-viable in a way SAM never was.
|
||
|
||
**It produces a flat partition with class ids**, so it populates `RegionField` directly: connected
|
||
components of the class map become regions, class boundaries become adjacency edges.
|
||
|
||
**What it does not produce is a hierarchy.** One partition at one semantic granularity. Click "sky"
|
||
and you get all the sky; there is no level at which you get *this part* of the sky. Adjacent
|
||
same-class regions merge whether or not you wanted them to — two different walls are one wall.
|
||
|
||
**Boundaries are semantically right and geometrically soft.** Internal stride is 4–8, upsampled to
|
||
H×W, so the class map is confident about *which* side of the boundary a pixel is on and vague about
|
||
*where* the boundary is to the pixel. Acceptable for biasing a contour. Not acceptable as a mask
|
||
edge at 100% zoom.
|
||
|
||
**Costs it brings that arm A does not:** a C dependency on the Android NDK against D13's policy, a
|
||
model to distribute and cache, an AGPL question (§7), an inference runtime per platform, and output
|
||
whose determinism across drivers is unproven (§6).
|
||
|
||
---
|
||
|
||
## 5. Arm C — semantic as a merge prior
|
||
|
||
The arms are not alternatives, and a comparison that omits their combination is a false dichotomy.
|
||
|
||
Weight each region-adjacency edge in arm A's union-find by boundary strength **and** by whether the
|
||
two regions share a semantic class. Regions that agree semantically merge earlier.
|
||
|
||
The result is a hierarchy whose coarse levels align with semantic objects and whose fine levels stay
|
||
pixel-accurate — the model doing what models are good at, which is knowing what things *are*, and
|
||
watershed doing what it is good at, which is knowing where boundaries are, exactly, at every scale.
|
||
It also repairs arm B's two weaknesses at once: the soft boundary is replaced by the watershed
|
||
boundary underneath it, and the missing granularity ladder is arm A's.
|
||
|
||
Arm C is the expected winner on quality. The question the spike actually has to answer is therefore
|
||
not "which is better" but **how much better than arm A alone, and is that increment worth D13's
|
||
cost.** §8 fixes that threshold in advance.
|
||
|
||
---
|
||
|
||
## 6. What gets measured
|
||
|
||
Per arm, over the corpus in §9, using the shared consumers from §2.
|
||
|
||
| # | Measure | Method | Why it decides anything |
|
||
|---|---|---|---|
|
||
| **M1** | **Interactions to target mask** | Clicks plus scroll steps to reach ≥95% IoU against a hand-traced mask | The real UX metric. "How many actions to get the mask I meant" is what a user experiences |
|
||
| **M2** | **Boundary accuracy** | Precision/recall of snapped-contour pixels within a 2px slack of the hand trace | Whether the edge survives 100% zoom, where masks are actually judged |
|
||
| **M3** | **Granularity coverage** | Per case, yes/no: does *any* hierarchy level produce the target region? | A hard failure mode. Arm B is expected to fail this wherever the target is not a class |
|
||
| **M4** | **Out-of-vocabulary behaviour** | M1 and M3 restricted to the OOV subset | Whether the arm degrades gracefully or produces nothing usable off-distribution |
|
||
| **M5** | **Determinism** | Same input twice on one machine; then across Mesa/AMD, NVIDIA, and Adreno | Gates whether a label field can be a cache key at all — see below |
|
||
| **M6** | **Precompute cost** | ms at proxy resolution and peak memory, on the reference desktop and one Android device | Whether it fits a background prefetch alongside the proxy |
|
||
| **M7** | **Distribution cost** | Added binary size, model size, new native dependencies, licence | D13's axis. Priced, not assumed |
|
||
|
||
**M5 deserves its own note, and it is a risk for arm A too.** ARCH §6.13 holds that cache keys are
|
||
computed over CPU-side *integer* state because GPU float results diverge across vendors. A label
|
||
field is integer state — but it is *derived from* float gradient arithmetic, so a boundary sitting
|
||
exactly between two basins could resolve differently on Adreno than on Mesa. If either arm proves
|
||
non-deterministic across vendors, its output cannot be a cache key and cannot round-trip through a
|
||
sidecar as region ids, which would push masks back toward rasters and undo most of §1's argument.
|
||
This is the measurement most likely to invalidate the whole approach, and it should be run early
|
||
rather than last.
|
||
|
||
---
|
||
|
||
## 7. Licence reading — before any code
|
||
|
||
D13's position applies unchanged: discovering at packaging time that a feature cannot ship is the
|
||
expensive failure, and it is entirely avoidable.
|
||
|
||
**Ultralytics ships YOLO under AGPL-3.0** — confirmed 2026-08-17 by reading `LICENSE` at the head of
|
||
`github.com/ultralytics/ultralytics`, which is the GNU Affero General Public License v3 verbatim.
|
||
That is deliberate on their part; the commercial licence is their business model.
|
||
|
||
GPLv3 §13 explicitly permits the combination, so this is *not* the blocker the InsightFace
|
||
non-commercial weights were: it is redistributable. But the combined work becomes effectively AGPL,
|
||
which is a change to DarkRoom's licensing posture rather than a dependency detail, and it needs to be
|
||
a decision made on purpose.
|
||
|
||
Still to verify before writing any of arm B:
|
||
|
||
- The licence on YOLO26 specifically, and on the ADE20K-pretrained weights *separately* from the
|
||
framework code — they are not necessarily the same grant.
|
||
- Whether ADE20K's own terms permit redistribution of weights derived from it.
|
||
- Whether AGPL is acceptable for DarkRoom, given Flatpak, F-Droid and Play distribution
|
||
(NFR-COMPAT-2).
|
||
|
||
**Arm A raises none of these questions**, which is worth stating plainly as part of its cost.
|
||
|
||
---
|
||
|
||
## 8. Decision criteria, fixed in advance
|
||
|
||
Stated now so the result cannot be rationalised afterwards.
|
||
|
||
- **Arm A ships alone** if it reaches within **one interaction** (M1) of arm C on the scene subset
|
||
*and* dominates arm C on the OOV subset (M4). The semantic increment does not then justify a C
|
||
dependency, an AGPL conversion, and a per-platform inference runtime.
|
||
- **Arm C ships** if it beats arm A by **two or more interactions** on the scene subset without
|
||
regressing OOV. That is a large enough difference to be felt in ordinary use, and it is what would
|
||
justify reopening D13.
|
||
- **Arm B never ships alone.** M3 is expected to fail on anything that is not an ADE20K class, and an
|
||
arm with a hard failure mode and no fallback is not a selection tool. If it surprises us and passes
|
||
M3 broadly, that is a genuine finding and this criterion is revisited on the evidence.
|
||
- **If M5 fails for an arm across vendors**, that arm cannot carry region ids into the sidecar
|
||
regardless of how it scored elsewhere.
|
||
|
||
---
|
||
|
||
## 9. Corpus
|
||
|
||
Roughly 24 images from a real library — three per category — hand-traced once and reused across all
|
||
arms. Categories chosen for the failure modes they provoke, not for coverage:
|
||
|
||
| Category | Provokes |
|
||
|---|---|
|
||
| Gradient sky | Low-contrast boundary; watershed banding |
|
||
| Foliage against sky | High-frequency boundary — arm A over-segments, arm B blurs |
|
||
| Hair against a busy background | The classic hard mask edge |
|
||
| Out-of-focus background | No edges at all; tests graceful failure in both |
|
||
| High-ISO noise | Arm A's known weakness; tests whether pre-smoothing is sufficient |
|
||
| Backlit silhouette | Strong unambiguous edge — the control case |
|
||
| Macro, abstract, still life | **OOV for ADE20K.** Arm B expected to fail M3 here |
|
||
| Architectural detail | Repeated structure; arm B merges distinct walls into one class |
|
||
|
||
Hand-tracing 24 masks is a couple of hours and it is what makes M1 and M2 mean anything. Without
|
||
ground truth this comparison is two demos and a preference.
|
||
|
||
---
|
||
|
||
## 10. Deliverables
|
||
|
||
Nothing in the UI, nothing in the graph, nothing in the sidecar.
|
||
|
||
- `core/dr-gpu/src/shaders/watershed.wgsl` — gradient and basin propagation.
|
||
- `core/dr-gpu/src/segment.rs` — the passes, producing a `RegionField`.
|
||
- RAG construction and the union-find merge tree as a pure-CPU module with unit tests and no device,
|
||
so the hierarchy is testable headless the way `dr-pipeline` is (ARCH §6.5a).
|
||
- The two shared consumers from §2 — live-wire over a summable cost field, and region-set selection.
|
||
- `core/dr-gpu/examples/segment.rs` — false-coloured PNGs at four or five hierarchy levels, plus the
|
||
M1/M2 numbers against the traced corpus.
|
||
|
||
It graduates to `core/dr-segment` if it ships; that is not a spike decision.
|
||
|
||
---
|
||
|
||
## 11. Order
|
||
|
||
1. **Licence reading (§7).** Hours, and it can eliminate arm B before anything is built.
|
||
2. **Arm A, and the shared consumers.** About a day. Look at the false-coloured PNGs — if the
|
||
granularity ladder does not feel right, nothing downstream matters and that is worth knowing
|
||
immediately.
|
||
3. **M5 across vendors, early.** It is the measurement that can invalidate the region-id
|
||
representation entirely, and it wants knowing before the corpus work is invested.
|
||
4. **The traced corpus, then M1–M4 on arm A.** Establishes the baseline every other arm is judged
|
||
against.
|
||
5. **Arms B and C**, only if §7 cleared and arm A's baseline leaves room worth closing.
|
||
|
||
Arm A is a day and needs no model, no runtime, no licence and no new dependency. It is also the
|
||
substrate every model-based arm writes into — so it is first regardless of how the comparison
|
||
eventually lands.
|
||
|
||
---
|
||
|
||
## 12. Arm A results
|
||
|
||
Built 2026-08-17. `core/dr-gpu/src/{segment.rs,hierarchy.rs}`,
|
||
`shaders/watershed.wgsl`, `examples/segment.rs`. 15 tests, 11 of them device-free.
|
||
|
||
**It works, and the hierarchy is not the expensive part.** On a 1200×800 synthetic at blur radius 2,
|
||
release build, RTX 3050 laptop: 6,730 basins and 19,223 boundaries found in **67 ms including the
|
||
readback**, and the merge tree built from them in **0.2 ms**. The tree is ~0.3% of the cost. The
|
||
estimate that priced it as a week's work was wrong by about two orders of magnitude, and the reason
|
||
is worth recording: it is Kruskal over a few thousand nodes, not a segmentation algorithm.
|
||
|
||
**The granularity ladder behaves.** At the fine end the background fragments badly — a smooth tonal
|
||
ramp bands into horizontal strips, and flat areas break into diagonal chains (see below). By
|
||
`cut_to(300)` all of that is gone: the hard-edged disc is exactly one region, the whole gradient
|
||
background is one region, and only genuine noise still fragments. The over-segmentation is absorbed
|
||
by the merge order rather than needing to be prevented, which is the property the whole design rests
|
||
on.
|
||
|
||
**Pre-smoothing is the knob it was claimed to be.** Radius 2 leaves the noisy corner fragmented at
|
||
300 regions; radius 5 largely clears it. Tying it to ISO is the right control.
|
||
|
||
Three findings that change what comes next:
|
||
|
||
**F1 — plateaux fragment into diagonal chains.** In an exactly flat region every pixel's steepest
|
||
descent is a tie, and the (value, index) tie-break sends them all up-left, so a plateau resolves into
|
||
diagonal streaks rather than one basin. Harmless here because those saddles are ~0 and the tree
|
||
merges them first — but a real sky or wall is a large plateau, and relying on the hierarchy to clean
|
||
up an artefact of the flow pass is fragile. The principled fix is a **lower-complete transform**: one
|
||
extra pass giving plateau pixels a gradient toward their nearest descending exit. Standard, cheap,
|
||
and worth doing before the corpus work.
|
||
|
||
*Attempted, and parked.* The pass exists — `plateau_init` seeds every pixel
|
||
that has a strictly lower neighbour, `plateau_step` carries a breadth-first
|
||
distance inward within a level set, and `flow` takes that distance as the
|
||
second key of a lexicographic tie-break. Bindings, ping-pong and dispatch were
|
||
all checked and are right. It is nonetheless a **measured no-op**: with a test
|
||
comparing the labelling at one iteration against sixty-four, *zero* of 9216
|
||
pixels change basin. That test is committed and ignored rather than deleted,
|
||
because it is the thing that turned "we think this works" into a fact.
|
||
|
||
Three explanations were tried and none of them was it. Exact float equality is
|
||
certainly wrong — a gradient computed from 8-bit samples is never exactly
|
||
equal across a region the eye calls flat — and a `LEVEL_EPS` tolerance now
|
||
replaces `==` and `<` in all three comparisons; it did not change the outcome.
|
||
Nor did the test image: a flat disc, a terraced disc and a constant-slope ramp
|
||
all behave identically. Worth knowing for whoever picks this up: on a
|
||
gradient-*magnitude* watershed, every flat region of the picture sits at
|
||
gradient zero, which is the global minimum, and a plateau with no descending
|
||
exit is a minimum — one basin by definition, with nothing for lower-completion
|
||
to resolve. The plateaux that do have an exit are regions of constant non-zero
|
||
gradient, which are rarer in a photograph than F1's phrasing suggests.
|
||
|
||
`plateau_iterations` therefore defaults to **0**. The pass is off, costs
|
||
nothing, and F1 stands open.
|
||
|
||
**F2 — `cut_to(N)` is a visualisation, not the interaction.** A global cut by region count spends its
|
||
budget wherever the saddles happen to be densest: at blur 5 the soft-edged disc's interior held a
|
||
cluster of near-equal saddles and ate the budget, fragmenting at a level where everything else was
|
||
clean. The real interaction walks up locally from the clicked region and has no such coupling. The
|
||
ladder in the example should not be read as what a user would experience.
|
||
|
||
**F3 — the RAG build still needs a readback.** `Segmentation::read_field` copies labels and gradient
|
||
to the CPU, gated behind the `readback` feature exactly as `read_pixels` is. Fine for a spike and
|
||
off the frame path, but a shipping build cannot take it (ARCH §6.1, AC-8), so the adjacency
|
||
accumulation has to move GPU-side with atomics. That is the largest known gap between this and
|
||
something shippable.
|
||
|
||
**M5 partially answered.** Run-to-run on one device is bit-identical, and the CPU half contributes no
|
||
nondeterminism of its own — both asserted by tests. Cross-vendor is untouched and remains the
|
||
measurement that can invalidate the region-id representation.
|
||
|
||
---
|
||
|
||
## 13. Arms B and C, and what §4 got wrong
|
||
|
||
Built 2026-08-21 on branch `local-adjustments`. `core/dr-segment`, `core/dr-gpu/src/mask.rs`,
|
||
`core/dr-pipeline/src/mask.rs`, and the develop panel.
|
||
|
||
**The dependency question dissolved rather than being decided.** §4 and D13 both priced arm B as
|
||
costing a C dependency under the Android NDK, and treated that as most of the difference between the
|
||
arms. It is not a cost that has to be paid: `ort` 2.0's `alternative-backend` feature disables the
|
||
linking entirely and lets another engine supply the `OrtApi`, and `ort-tract` — same authors,
|
||
MIT/Apache — supplies it from `tract`, which is pure Rust. So arm B runs through `ort`'s API with no
|
||
C anywhere, and D13's "largest exception the policy would tolerate" turns out not to be needed.
|
||
|
||
Measured before committing to it, because tract's operator coverage is the thing that could have
|
||
sunk it: **yolo26n-seg loads with zero unsupported operators** and runs 640×640 in ~470 ms on the
|
||
reference desktop's CPU. Correct masks on the standard `bus.jpg` — one bus and three people, outlines
|
||
following the subjects.
|
||
|
||
Three findings that contradict §4 directly, and all three change the design rather than the schedule.
|
||
|
||
**F4 — there is no ADE20K-trained YOLO.** §4's whole argument for arm B was ADE20K's 150 classes and
|
||
their *stuff* categories: "'that patch of sky' is a class here." Checked 2026-08-21: Ultralytics ships
|
||
YOLO26-seg trained on **COCO**, whose 80 classes are all *things*, and the one HuggingFace repository
|
||
claiming a YOLO/ADE20K combination is empty. ADE20K models exist as SegFormer/OneFormer/MaskFormer
|
||
transformers, not as YOLO.
|
||
|
||
So the shipped vocabulary selects **subjects**, not **stuff**. "Select the person" works; "select the
|
||
sky" does not come from the model at all and must come from the watershed. That is a narrower arm B
|
||
than §4 assumed, and it *raises* the importance of arm C rather than lowering it — the model can no
|
||
longer be the whole answer for anything.
|
||
|
||
**F5 — it is instance segmentation, not semantic segmentation.** §4 assumed a flat partition with
|
||
class ids that would "populate `RegionField` directly". YOLO-seg does not partition the image; it
|
||
finds objects, and most pixels in a landscape belong to no instance. Two consequences, one bad and
|
||
one better than expected: nothing populates a `RegionField` on its own, and two people come back as
|
||
*two* instances where a semantic model would have returned one "person" area covering both. For
|
||
selecting a subject the latter is the behaviour worth having.
|
||
|
||
**F6 — tract cannot parse a dynamic-shape export.** It fails shape inference on the neck's `Concat`.
|
||
The graph therefore ships with its input fixed at 640×640 square, and every image is letterboxed into
|
||
it. This is the constraint behind the tiling option in `semantic.rs`: with a fixed window, tiling is
|
||
the *only* route to more semantic resolution, and it costs one inference per tile (≈2.8 s for a 3×2
|
||
grid over a 1600 px proxy against 470 ms whole-frame). Defaulted off — a photographic subject is
|
||
usually large in frame, which is the case whole-frame inference handles best — and left implemented
|
||
so §9's corpus can settle it rather than an argument.
|
||
|
||
**Arm C ships, and the §8 criteria were not what decided it.** §8 asked for a two-interaction margin
|
||
over arm A on the scene subset. That comparison was never run, because F4 and F5 changed what the
|
||
arms *are*: with a model that recognises subjects and has no word for sky, arm B alone cannot be a
|
||
selection tool at all (§8's "arm B never ships alone" holds, for a stronger reason than expected),
|
||
and arm A alone cannot tell a person from the wall behind them. They are complements rather than
|
||
candidates. Arm C's implementation is `prior.rs`: instance membership re-weights the merge saddles,
|
||
so region pairs the model believes share an object merge early and pairs straddling its edge merge
|
||
late. **No boundary moves** — only the order in which boundaries dissolve — which is how the result
|
||
stays pixel-accurate at every level while its coarse levels become named things.
|
||
|
||
**M1–M4 remain unmeasured.** The 24-image corpus of §9 has not been traced, so there are no
|
||
interaction counts and no boundary-accuracy numbers. What exists is a working feature and the
|
||
evidence that each piece does what it claims in isolation. The corpus is still the thing that would
|
||
turn "this feels right" into a number, and it is the largest piece of §11 left undone.
|
||
|
||
**M5 is unchanged and still the risk it was.** Run-to-run on one device is identical, asserted by a
|
||
test. Cross-vendor is untouched. Region ids now reach the sidecar, so if the label field proves
|
||
non-deterministic across vendors a mask written on the desktop will not mean the same thing on
|
||
Android — see `MaskSource::Regions::signature`, which detects a *retuned* segmentation but not a
|
||
differently-rounded one.
|
||
|
||
**F3 still stands.** `Segmentation::read_field` still copies the label and gradient buffers to the
|
||
CPU to build the region graph. It is now behind its own `segment-readback` feature rather than
|
||
sharing `readback` — this transfer is once per image on a worker, where the one AC-8 forbids is per
|
||
frame in the render loop — but the accumulation still belongs GPU-side with atomics.
|
||
|
||
---
|
||
|
||
## 14. Register entries
|
||
|
||
**S15** — *Region segmentation for local masking* · **CLOSED 2026-08-21**. Arms A, B and C built; the
|
||
shared consumers built; the licence question resolved (§7, D14). The 24-image corpus was not traced,
|
||
so M1–M4 are unmeasured and M5 is answered only on one device. Answers: local masking snaps to arm C,
|
||
and a mask is stored as region ids. Relates to: D13, D14, FR-DEV-3, ARCH §5.4, §6.13.
|
||
|
||
**D14** — *Segmentation source for local masking* · **DECIDED 2026-08-21**: **arm C**, a watershed
|
||
hierarchy re-weighted by YOLO26n-seg instance membership, with the model optional and the watershed
|
||
sufficient without it. Weights ship in-tree under AGPL-3.0, which GPLv3 §13 permits and which makes
|
||
the combined work effectively AGPL — a deliberate change to DarkRoom's licensing posture, not a
|
||
dependency detail (`core/dr-segment/models/LICENCE.md`).
|
||
|
||
**D13** — *inference runtime* · the dependency half is **answered** for segmentation and the answer
|
||
generalises: `ort` + `ort-tract` gives ONNX inference in pure Rust, so the face pipeline of §3.9.1
|
||
needs no C dependency either. The *model licensing* half of D13 is untouched — the InsightFace
|
||
weights are still non-commercial and still unusable here.
|
||
|
||
---
|
||
|
||
## 16. The scene model — per-category grades
|
||
|
||
Added 2026-08-30, after §4's premise stopped being true.
|
||
|
||
### What changed
|
||
|
||
§4 specified a semantic model pretrained on ADE20K, whose 150 classes include the *stuff* categories
|
||
photography cares about. §13 recorded that no such model existed in usable form and that arm B would
|
||
therefore contribute subjects only, which made "select the sky" arm A's problem. Re-checked
|
||
2026-08-30: **Ultralytics now ships a `semantic` task with ADE20K checkpoints**
|
||
(`docs.ultralytics.com/tasks/semantic`). `yolo26s-sem-ade20k` is in `models/scene/`.
|
||
|
||
### It is an addition, not a correction to arm B
|
||
|
||
The instance model stays exactly where it was, and the reason is the one §13 already gave and was
|
||
right about: a semantic model merges every pixel of a class into one region, so it cannot separate
|
||
two people, and separating two people is what clicking a subject requires. Swapping arm B for this
|
||
would regress the primary interaction to fix a secondary one.
|
||
|
||
So the two divide by *what the user is doing*, not by which is better:
|
||
|
||
| | `models/segment/` (COCO instances) | `models/scene/` (ADE20K semantics) |
|
||
|---|---|---|
|
||
| Question | which pixels are *that* dog | how much of this pixel is sky |
|
||
| Granularity | per instance | per category, whole frame |
|
||
| Drives | local adjustments, subject selection | the scene tab's per-category sliders |
|
||
| Vocabulary | 80 things | 150 classes, stuff included |
|
||
|
||
### The export is truncated, and both reasons matter
|
||
|
||
Ultralytics ends the graph with `Resize → ArgMax → Cast`, returning a `[1, 640, 640]` u8 label map.
|
||
`tools/export-seg-model.sh` cuts that tail and ships the classifier's `[1, 150, 80, 80]` f32 logits.
|
||
|
||
**Cost.** The `Resize` materialises 150 × 640 × 640 × f32 — 246 MB — and the `ArgMax` then reduces
|
||
across the channel axis, striding 409,600 elements per comparison. Measured under load it was
|
||
roughly four fifths of total runtime, spent on work the application discards.
|
||
|
||
**Softness, which is the more important one.** `ArgMax` destroys the per-class scores, and the whole
|
||
design of the scene tab rests on keeping them. Softmax over the 150 channels, summed within each
|
||
category, produces per-category weights that sum to one at every pixel — a partition of unity.
|
||
Feathering that cannot double-grade a boundary. Feathering *hard labels* outward from two adjacent
|
||
categories paints both grades into the overlap, and every horizon in the frame acquires a seam.
|
||
|
||
### The resolution is 80×80, and no setting changes that
|
||
|
||
The discarded upsample was never information. `Scene` keeps the native grid and resamples on demand,
|
||
so the coarseness is visible in the type rather than hidden. A graduated grade over sky or water is
|
||
untroubled by it; a rooftop against sky at 100% zoom will show it. This is the constraint most likely
|
||
to decide whether the tab feels good, and it is not addressable by choosing a larger checkpoint —
|
||
`yolo26n-sem` and `yolo26s-sem` have the same output grid.
|
||
|
||
### Licence
|
||
|
||
Unchanged. Same AGPL-3.0 grant as the instance model, same GPLv3 §13 permission, same consequence
|
||
already accepted in D14 — so this needed no new licence decision, which is most of why it was cheap.
|
||
See `models/LICENCE.md`.
|
||
|
||
### Measurement
|
||
|
||
Timings taken while this was chosen came off a laptop compiling other things and are upper bounds
|
||
only. `cargo run -p dr-segment --example scene --release --features embedded-scene-model` reports a
|
||
median over N runs with the first excluded; a number worth quoting should come from that, on an idle
|
||
machine.
|