Files
DarkRoom/docs/dev/segmentation.md
dtourolle 84fade99ec Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.

Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
2026-09-20 21:16:03 +02:00

504 lines
29 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Region segmentation for local masking
Spec for **S15**, the spike that decides how DarkRoom finds the boundaries a local mask snaps to.
Local adjustments (FR-DEV-3, "linear gradient, radial gradient, and brush masks") need more than
placement handles to be competitive. The interactions that matter — click to select a region, drag a
contour that clings to an edge, paint without crossing a boundary — all need the same thing
underneath: **a map of where the image's regions are.**
There are two credible ways to produce that map and they are not obviously ordered. This document
specifies both, specifies the third option of combining them, and fixes the measurements that decide
between them *before* any of them is built.
---
## 1. Why this is a spike and not a build
Three properties make the choice expensive to get wrong.
**It sets the mask representation.** If regions exist, a mask is a *set of region ids* — integers,
diffable, mergeable at node level under FR-NC-9, cheap in a sidecar. If they don't, a mask is a
raster, and rasters are none of those things. This is the decision that is expensive to retrofit;
everything else in local masking sits on top of it.
**One arm collides with a settled policy.** D13 records that every dependency choice in this project
has gone the same way — rustls over aws-lc-rs, bundled SQLite, a Rust Lensfun port, zune-jpeg — to
avoid a C dependency under the Android NDK, and names `ort` as the largest exception that policy
would tolerate. Arm B needs exactly that exception. Arm A needs no new dependency at all. That
asymmetry is not a tiebreak, it is most of the cost difference, and it should be priced honestly
rather than discovered at packaging time.
**Model licensing is a distribution blocker.** D13 already establishes this for the face pipeline,
and the same reading applies here — see §7. It is a licence-reading exercise, not a research
question, and it comes first.
---
## 2. The common interface
Both arms produce the same thing. This is what makes them comparable, and what lets the choice be
deferred behind a seam rather than baked into every consumer.
```rust
/// A partition of the image into labelled regions.
pub struct RegionField {
/// Per-pixel region id at proxy resolution. R32Uint on the GPU.
labels: Texture,
/// Per-region summary: pixel count, bounding box, mean colour,
/// and (arm B only) a semantic class id.
regions: Vec<Region>,
/// Boundary strength per adjacent region pair — the edge weight
/// the merge tree is built from and the cost field reads.
adjacency: Vec<(RegionId, RegionId, f32)>,
}
```
Two consumers sit on it, and neither knows which arm produced it:
**A cost field, for contour snapping.** Live-wire — Dijkstra from the last anchor to the cursor over
a per-pixel cost that is *low* on boundaries. The cost is a **sum of terms**, which is the property
that matters: image gradient is always available, region boundary strength is added when a
`RegionField` exists, and a semantic boundary term is added when a model is present. Each source
improves the snap without changing the interface, so the arms are not exclusive here even in
principle.
**A region set, for click selection.** Click reads the label under the cursor; the mask is
`label(px) ∈ selected`. Add and subtract are set operations on ids. No flood fill, no readback, no
iteration — the whole reason the precomputed map is worth having.
Both consumers are built once, in the spike, and shared by both arms. A comparison in which each arm
gets its own consumer measures the consumers.
---
## 3. Arm A — multiscale watershed
No model, no new dependency, deterministic, works on any image.
**Gradient.** Sobel magnitude over a perceptual luma plus chroma distance, not camera-space RGB —
channel-weighted RGB gradient reads a saturated red edge as weaker than it looks. Computed after
demosaic and denoise, before the edit graph, so an exposure change does not invalidate it.
**Pre-smoothing is not optional.** Raw watershed on a noisy file makes every grain its own basin.
A guided or bilateral pre-filter, with strength tied to the file's ISO, is part of the arm rather
than a refinement of it.
**Basins.** Each pixel points downhill to its steepest neighbour; pointer-jumping resolves every
pixel to its basin root in log passes. Two compute shaders and a dispatch loop.
**The hierarchy is the cheap part.** Build the region adjacency graph, sort edges by boundary
strength, union-find over them, and *record the merge order*. That recording is the merge tree — a
click selects a leaf, and a scroll walks up through progressively coarser merges. Textbook Kruskal on
a graph of a few thousand nodes.
That node count is why the tree build is a legitimate CPU operation: it runs on the adjacency graph,
not on pixels. Pixels stay on the GPU, the graph is CPU-side — the same split ARCH §3.4 and §6.1
already draw for the edit graph, so no exception to the no-readback rule is needed.
**Known weaknesses, to be measured rather than argued about:** over-segmentation on noise and
texture, weak boundaries where contrast is low but semantics are obvious (a pale sky meeting a pale
wall), and a granularity ladder that is geometric rather than meaningful — level 7 is *a* coarser
partition, not necessarily *the* object.
---
## 4. Arm B — semantic segmentation
YOLO26-seg pretrained on ADE20K, run through `ort`.
ADE20K's 150 classes include stuff — sky, vegetation, water, wall, road — which is a far better
vocabulary for photography than COCO's 80 thing-classes. "That patch of sky" is a class here. The
nano variant is ~1.6M parameters, which is genuinely mobile-viable in a way SAM never was.
**It produces a flat partition with class ids**, so it populates `RegionField` directly: connected
components of the class map become regions, class boundaries become adjacency edges.
**What it does not produce is a hierarchy.** One partition at one semantic granularity. Click "sky"
and you get all the sky; there is no level at which you get *this part* of the sky. Adjacent
same-class regions merge whether or not you wanted them to — two different walls are one wall.
**Boundaries are semantically right and geometrically soft.** Internal stride is 4–8, upsampled to
H×W, so the class map is confident about *which* side of the boundary a pixel is on and vague about
*where* the boundary is to the pixel. Acceptable for biasing a contour. Not acceptable as a mask
edge at 100% zoom.
**Costs it brings that arm A does not:** a C dependency on the Android NDK against D13's policy, a
model to distribute and cache, an AGPL question (§7), an inference runtime per platform, and output
whose determinism across drivers is unproven (§6).
---
## 5. Arm C — semantic as a merge prior
The arms are not alternatives, and a comparison that omits their combination is a false dichotomy.
Weight each region-adjacency edge in arm A's union-find by boundary strength **and** by whether the
two regions share a semantic class. Regions that agree semantically merge earlier.
The result is a hierarchy whose coarse levels align with semantic objects and whose fine levels stay
pixel-accurate — the model doing what models are good at, which is knowing what things *are*, and
watershed doing what it is good at, which is knowing where boundaries are, exactly, at every scale.
It also repairs arm B's two weaknesses at once: the soft boundary is replaced by the watershed
boundary underneath it, and the missing granularity ladder is arm A's.
Arm C is the expected winner on quality. The question the spike actually has to answer is therefore
not "which is better" but **how much better than arm A alone, and is that increment worth D13's
cost.** §8 fixes that threshold in advance.
---
## 6. What gets measured
Per arm, over the corpus in §9, using the shared consumers from §2.
| # | Measure | Method | Why it decides anything |
|---|---|---|---|
| **M1** | **Interactions to target mask** | Clicks plus scroll steps to reach ≥95% IoU against a hand-traced mask | The real UX metric. "How many actions to get the mask I meant" is what a user experiences |
| **M2** | **Boundary accuracy** | Precision/recall of snapped-contour pixels within a 2px slack of the hand trace | Whether the edge survives 100% zoom, where masks are actually judged |
| **M3** | **Granularity coverage** | Per case, yes/no: does *any* hierarchy level produce the target region? | A hard failure mode. Arm B is expected to fail this wherever the target is not a class |
| **M4** | **Out-of-vocabulary behaviour** | M1 and M3 restricted to the OOV subset | Whether the arm degrades gracefully or produces nothing usable off-distribution |
| **M5** | **Determinism** | Same input twice on one machine; then across Mesa/AMD, NVIDIA, and Adreno | Gates whether a label field can be a cache key at all — see below |
| **M6** | **Precompute cost** | ms at proxy resolution and peak memory, on the reference desktop and one Android device | Whether it fits a background prefetch alongside the proxy |
| **M7** | **Distribution cost** | Added binary size, model size, new native dependencies, licence | D13's axis. Priced, not assumed |
**M5 deserves its own note, and it is a risk for arm A too.** ARCH §6.13 holds that cache keys are
computed over CPU-side *integer* state because GPU float results diverge across vendors. A label
field is integer state — but it is *derived from* float gradient arithmetic, so a boundary sitting
exactly between two basins could resolve differently on Adreno than on Mesa. If either arm proves
non-deterministic across vendors, its output cannot be a cache key and cannot round-trip through a
sidecar as region ids, which would push masks back toward rasters and undo most of §1's argument.
This is the measurement most likely to invalidate the whole approach, and it should be run early
rather than last.
---
## 7. Licence reading — before any code
D13's position applies unchanged: discovering at packaging time that a feature cannot ship is the
expensive failure, and it is entirely avoidable.
**Ultralytics ships YOLO under AGPL-3.0** — confirmed 2026-08-17 by reading `LICENSE` at the head of
`github.com/ultralytics/ultralytics`, which is the GNU Affero General Public License v3 verbatim.
That is deliberate on their part; the commercial licence is their business model.
GPLv3 §13 explicitly permits the combination, so this is *not* the blocker the InsightFace
non-commercial weights were: it is redistributable. But the combined work becomes effectively AGPL,
which is a change to DarkRoom's licensing posture rather than a dependency detail, and it needs to be
a decision made on purpose.
Still to verify before writing any of arm B:
- The licence on YOLO26 specifically, and on the ADE20K-pretrained weights *separately* from the
framework code — they are not necessarily the same grant.
- Whether ADE20K's own terms permit redistribution of weights derived from it.
- Whether AGPL is acceptable for DarkRoom, given Flatpak, F-Droid and Play distribution
(NFR-COMPAT-2).
**Arm A raises none of these questions**, which is worth stating plainly as part of its cost.
---
## 8. Decision criteria, fixed in advance
Stated now so the result cannot be rationalised afterwards.
- **Arm A ships alone** if it reaches within **one interaction** (M1) of arm C on the scene subset
*and* dominates arm C on the OOV subset (M4). The semantic increment does not then justify a C
dependency, an AGPL conversion, and a per-platform inference runtime.
- **Arm C ships** if it beats arm A by **two or more interactions** on the scene subset without
regressing OOV. That is a large enough difference to be felt in ordinary use, and it is what would
justify reopening D13.
- **Arm B never ships alone.** M3 is expected to fail on anything that is not an ADE20K class, and an
arm with a hard failure mode and no fallback is not a selection tool. If it surprises us and passes
M3 broadly, that is a genuine finding and this criterion is revisited on the evidence.
- **If M5 fails for an arm across vendors**, that arm cannot carry region ids into the sidecar
regardless of how it scored elsewhere.
---
## 9. Corpus
Roughly 24 images from a real library — three per category — hand-traced once and reused across all
arms. Categories chosen for the failure modes they provoke, not for coverage:
| Category | Provokes |
|---|---|
| Gradient sky | Low-contrast boundary; watershed banding |
| Foliage against sky | High-frequency boundary — arm A over-segments, arm B blurs |
| Hair against a busy background | The classic hard mask edge |
| Out-of-focus background | No edges at all; tests graceful failure in both |
| High-ISO noise | Arm A's known weakness; tests whether pre-smoothing is sufficient |
| Backlit silhouette | Strong unambiguous edge — the control case |
| Macro, abstract, still life | **OOV for ADE20K.** Arm B expected to fail M3 here |
| Architectural detail | Repeated structure; arm B merges distinct walls into one class |
Hand-tracing 24 masks is a couple of hours and it is what makes M1 and M2 mean anything. Without
ground truth this comparison is two demos and a preference.
---
## 10. Deliverables
Nothing in the UI, nothing in the graph, nothing in the sidecar.
- `core/dr-gpu/src/shaders/watershed.wgsl` — gradient and basin propagation.
- `core/dr-gpu/src/segment.rs` — the passes, producing a `RegionField`.
- RAG construction and the union-find merge tree as a pure-CPU module with unit tests and no device,
so the hierarchy is testable headless the way `dr-pipeline` is (ARCH §6.5a).
- The two shared consumers from §2 — live-wire over a summable cost field, and region-set selection.
- `core/dr-gpu/examples/segment.rs` — false-coloured PNGs at four or five hierarchy levels, plus the
M1/M2 numbers against the traced corpus.
It graduates to `core/dr-segment` if it ships; that is not a spike decision.
---
## 11. Order
1. **Licence reading (§7).** Hours, and it can eliminate arm B before anything is built.
2. **Arm A, and the shared consumers.** About a day. Look at the false-coloured PNGs — if the
granularity ladder does not feel right, nothing downstream matters and that is worth knowing
immediately.
3. **M5 across vendors, early.** It is the measurement that can invalidate the region-id
representation entirely, and it wants knowing before the corpus work is invested.
4. **The traced corpus, then M1–M4 on arm A.** Establishes the baseline every other arm is judged
against.
5. **Arms B and C**, only if §7 cleared and arm A's baseline leaves room worth closing.
Arm A is a day and needs no model, no runtime, no licence and no new dependency. It is also the
substrate every model-based arm writes into — so it is first regardless of how the comparison
eventually lands.
---
## 12. Arm A results
Built 2026-08-17. `core/dr-gpu/src/{segment.rs,hierarchy.rs}`,
`shaders/watershed.wgsl`, `examples/segment.rs`. 15 tests, 11 of them device-free.
**It works, and the hierarchy is not the expensive part.** On a 1200×800 synthetic at blur radius 2,
release build, RTX 3050 laptop: 6,730 basins and 19,223 boundaries found in **67 ms including the
readback**, and the merge tree built from them in **0.2 ms**. The tree is ~0.3% of the cost. The
estimate that priced it as a week's work was wrong by about two orders of magnitude, and the reason
is worth recording: it is Kruskal over a few thousand nodes, not a segmentation algorithm.
**The granularity ladder behaves.** At the fine end the background fragments badly — a smooth tonal
ramp bands into horizontal strips, and flat areas break into diagonal chains (see below). By
`cut_to(300)` all of that is gone: the hard-edged disc is exactly one region, the whole gradient
background is one region, and only genuine noise still fragments. The over-segmentation is absorbed
by the merge order rather than needing to be prevented, which is the property the whole design rests
on.
**Pre-smoothing is the knob it was claimed to be.** Radius 2 leaves the noisy corner fragmented at
300 regions; radius 5 largely clears it. Tying it to ISO is the right control.
Three findings that change what comes next:
**F1 — plateaux fragment into diagonal chains.** In an exactly flat region every pixel's steepest
descent is a tie, and the (value, index) tie-break sends them all up-left, so a plateau resolves into
diagonal streaks rather than one basin. Harmless here because those saddles are ~0 and the tree
merges them first — but a real sky or wall is a large plateau, and relying on the hierarchy to clean
up an artefact of the flow pass is fragile. The principled fix is a **lower-complete transform**: one
extra pass giving plateau pixels a gradient toward their nearest descending exit. Standard, cheap,
and worth doing before the corpus work.
*Attempted, and parked.* The pass exists — `plateau_init` seeds every pixel
that has a strictly lower neighbour, `plateau_step` carries a breadth-first
distance inward within a level set, and `flow` takes that distance as the
second key of a lexicographic tie-break. Bindings, ping-pong and dispatch were
all checked and are right. It is nonetheless a **measured no-op**: with a test
comparing the labelling at one iteration against sixty-four, *zero* of 9216
pixels change basin. That test is committed and ignored rather than deleted,
because it is the thing that turned "we think this works" into a fact.
Three explanations were tried and none of them was it. Exact float equality is
certainly wrong — a gradient computed from 8-bit samples is never exactly
equal across a region the eye calls flat — and a `LEVEL_EPS` tolerance now
replaces `==` and `<` in all three comparisons; it did not change the outcome.
Nor did the test image: a flat disc, a terraced disc and a constant-slope ramp
all behave identically. Worth knowing for whoever picks this up: on a
gradient-*magnitude* watershed, every flat region of the picture sits at
gradient zero, which is the global minimum, and a plateau with no descending
exit is a minimum — one basin by definition, with nothing for lower-completion
to resolve. The plateaux that do have an exit are regions of constant non-zero
gradient, which are rarer in a photograph than F1's phrasing suggests.
`plateau_iterations` therefore defaults to **0**. The pass is off, costs
nothing, and F1 stands open.
**F2 — `cut_to(N)` is a visualisation, not the interaction.** A global cut by region count spends its
budget wherever the saddles happen to be densest: at blur 5 the soft-edged disc's interior held a
cluster of near-equal saddles and ate the budget, fragmenting at a level where everything else was
clean. The real interaction walks up locally from the clicked region and has no such coupling. The
ladder in the example should not be read as what a user would experience.
**F3 — the RAG build still needs a readback.** `Segmentation::read_field` copies labels and gradient
to the CPU, gated behind the `readback` feature exactly as `read_pixels` is. Fine for a spike and
off the frame path, but a shipping build cannot take it (ARCH §6.1, AC-8), so the adjacency
accumulation has to move GPU-side with atomics. That is the largest known gap between this and
something shippable.
**M5 partially answered.** Run-to-run on one device is bit-identical, and the CPU half contributes no
nondeterminism of its own — both asserted by tests. Cross-vendor is untouched and remains the
measurement that can invalidate the region-id representation.
---
## 13. Arms B and C, and what §4 got wrong
Built 2026-08-21 on branch `local-adjustments`. `core/dr-segment`, `core/dr-gpu/src/mask.rs`,
`core/dr-pipeline/src/mask.rs`, and the develop panel.
**The dependency question dissolved rather than being decided.** §4 and D13 both priced arm B as
costing a C dependency under the Android NDK, and treated that as most of the difference between the
arms. It is not a cost that has to be paid: `ort` 2.0's `alternative-backend` feature disables the
linking entirely and lets another engine supply the `OrtApi`, and `ort-tract` — same authors,
MIT/Apache — supplies it from `tract`, which is pure Rust. So arm B runs through `ort`'s API with no
C anywhere, and D13's "largest exception the policy would tolerate" turns out not to be needed.
Measured before committing to it, because tract's operator coverage is the thing that could have
sunk it: **yolo26n-seg loads with zero unsupported operators** and runs 640×640 in ~470 ms on the
reference desktop's CPU. Correct masks on the standard `bus.jpg` — one bus and three people, outlines
following the subjects.
Three findings that contradict §4 directly, and all three change the design rather than the schedule.
**F4 — there is no ADE20K-trained YOLO.** §4's whole argument for arm B was ADE20K's 150 classes and
their *stuff* categories: "'that patch of sky' is a class here." Checked 2026-08-21: Ultralytics ships
YOLO26-seg trained on **COCO**, whose 80 classes are all *things*, and the one HuggingFace repository
claiming a YOLO/ADE20K combination is empty. ADE20K models exist as SegFormer/OneFormer/MaskFormer
transformers, not as YOLO.
So the shipped vocabulary selects **subjects**, not **stuff**. "Select the person" works; "select the
sky" does not come from the model at all and must come from the watershed. That is a narrower arm B
than §4 assumed, and it *raises* the importance of arm C rather than lowering it — the model can no
longer be the whole answer for anything.
**F5 — it is instance segmentation, not semantic segmentation.** §4 assumed a flat partition with
class ids that would "populate `RegionField` directly". YOLO-seg does not partition the image; it
finds objects, and most pixels in a landscape belong to no instance. Two consequences, one bad and
one better than expected: nothing populates a `RegionField` on its own, and two people come back as
*two* instances where a semantic model would have returned one "person" area covering both. For
selecting a subject the latter is the behaviour worth having.
**F6 — tract cannot parse a dynamic-shape export.** It fails shape inference on the neck's `Concat`.
The graph therefore ships with its input fixed at 640×640 square, and every image is letterboxed into
it. This is the constraint behind the tiling option in `semantic.rs`: with a fixed window, tiling is
the *only* route to more semantic resolution, and it costs one inference per tile (≈2.8 s for a 3×2
grid over a 1600 px proxy against 470 ms whole-frame). Defaulted off — a photographic subject is
usually large in frame, which is the case whole-frame inference handles best — and left implemented
so §9's corpus can settle it rather than an argument.
**Arm C ships, and the §8 criteria were not what decided it.** §8 asked for a two-interaction margin
over arm A on the scene subset. That comparison was never run, because F4 and F5 changed what the
arms *are*: with a model that recognises subjects and has no word for sky, arm B alone cannot be a
selection tool at all (§8's "arm B never ships alone" holds, for a stronger reason than expected),
and arm A alone cannot tell a person from the wall behind them. They are complements rather than
candidates. Arm C's implementation is `prior.rs`: instance membership re-weights the merge saddles,
so region pairs the model believes share an object merge early and pairs straddling its edge merge
late. **No boundary moves** — only the order in which boundaries dissolve — which is how the result
stays pixel-accurate at every level while its coarse levels become named things.
**M1–M4 remain unmeasured.** The 24-image corpus of §9 has not been traced, so there are no
interaction counts and no boundary-accuracy numbers. What exists is a working feature and the
evidence that each piece does what it claims in isolation. The corpus is still the thing that would
turn "this feels right" into a number, and it is the largest piece of §11 left undone.
**M5 is unchanged and still the risk it was.** Run-to-run on one device is identical, asserted by a
test. Cross-vendor is untouched. Region ids now reach the sidecar, so if the label field proves
non-deterministic across vendors a mask written on the desktop will not mean the same thing on
Android — see `MaskSource::Regions::signature`, which detects a *retuned* segmentation but not a
differently-rounded one.
**F3 still stands.** `Segmentation::read_field` still copies the label and gradient buffers to the
CPU to build the region graph. It is now behind its own `segment-readback` feature rather than
sharing `readback` — this transfer is once per image on a worker, where the one AC-8 forbids is per
frame in the render loop — but the accumulation still belongs GPU-side with atomics.
---
## 14. Register entries
**S15** — *Region segmentation for local masking* · **CLOSED 2026-08-21**. Arms A, B and C built; the
shared consumers built; the licence question resolved (§7, D14). The 24-image corpus was not traced,
so M1–M4 are unmeasured and M5 is answered only on one device. Answers: local masking snaps to arm C,
and a mask is stored as region ids. Relates to: D13, D14, FR-DEV-3, ARCH §5.4, §6.13.
**D14** — *Segmentation source for local masking* · **DECIDED 2026-08-21**: **arm C**, a watershed
hierarchy re-weighted by YOLO26n-seg instance membership, with the model optional and the watershed
sufficient without it. Weights ship in-tree under AGPL-3.0, which GPLv3 §13 permits and which makes
the combined work effectively AGPL — a deliberate change to DarkRoom's licensing posture, not a
dependency detail (`core/dr-segment/models/LICENCE.md`).
**D13** — *inference runtime* · the dependency half is **answered** for segmentation and the answer
generalises: `ort` + `ort-tract` gives ONNX inference in pure Rust, so the face pipeline of §3.9.1
needs no C dependency either. The *model licensing* half of D13 is untouched — the InsightFace
weights are still non-commercial and still unusable here.
---
## 16. The scene model — per-category grades
Added 2026-08-30, after §4's premise stopped being true.
### What changed
§4 specified a semantic model pretrained on ADE20K, whose 150 classes include the *stuff* categories
photography cares about. §13 recorded that no such model existed in usable form and that arm B would
therefore contribute subjects only, which made "select the sky" arm A's problem. Re-checked
2026-08-30: **Ultralytics now ships a `semantic` task with ADE20K checkpoints**
(`docs.ultralytics.com/tasks/semantic`). `yolo26s-sem-ade20k` is in `models/scene/`.
### It is an addition, not a correction to arm B
The instance model stays exactly where it was, and the reason is the one §13 already gave and was
right about: a semantic model merges every pixel of a class into one region, so it cannot separate
two people, and separating two people is what clicking a subject requires. Swapping arm B for this
would regress the primary interaction to fix a secondary one.
So the two divide by *what the user is doing*, not by which is better:
| | `models/segment/` (COCO instances) | `models/scene/` (ADE20K semantics) |
|---|---|---|
| Question | which pixels are *that* dog | how much of this pixel is sky |
| Granularity | per instance | per category, whole frame |
| Drives | local adjustments, subject selection | the scene tab's per-category sliders |
| Vocabulary | 80 things | 150 classes, stuff included |
### The export is truncated, and both reasons matter
Ultralytics ends the graph with `Resize → ArgMax → Cast`, returning a `[1, 640, 640]` u8 label map.
`tools/export-seg-model.sh` cuts that tail and ships the classifier's `[1, 150, 80, 80]` f32 logits.
**Cost.** The `Resize` materialises 150 × 640 × 640 × f32 — 246 MB — and the `ArgMax` then reduces
across the channel axis, striding 409,600 elements per comparison. Measured under load it was
roughly four fifths of total runtime, spent on work the application discards.
**Softness, which is the more important one.** `ArgMax` destroys the per-class scores, and the whole
design of the scene tab rests on keeping them. Softmax over the 150 channels, summed within each
category, produces per-category weights that sum to one at every pixel — a partition of unity.
Feathering that cannot double-grade a boundary. Feathering *hard labels* outward from two adjacent
categories paints both grades into the overlap, and every horizon in the frame acquires a seam.
### The resolution is 80×80, and no setting changes that
The discarded upsample was never information. `Scene` keeps the native grid and resamples on demand,
so the coarseness is visible in the type rather than hidden. A graduated grade over sky or water is
untroubled by it; a rooftop against sky at 100% zoom will show it. This is the constraint most likely
to decide whether the tab feels good, and it is not addressable by choosing a larger checkpoint —
`yolo26n-sem` and `yolo26s-sem` have the same output grid.
### Licence
Unchanged. Same AGPL-3.0 grant as the instance model, same GPLv3 §13 permission, same consequence
already accepted in D14 — so this needed no new licence decision, which is most of why it was cheap.
See `models/LICENCE.md`.
### Measurement
Timings taken while this was chosen came off a laptop compiling other things and are upper bounds
only. `cargo run -p dr-segment --example scene --release --features embedded-scene-model` reports a
median over N runs with the first excluded; a number worth quoting should come from that, on an idle
machine.