Picks up the lower-completion work a crashed session left mid-debug, with one failing test and no diagnosis. The diagnosis is that the pass is a no-op. Not "does not reduce basin count" — it changes *no pixel's basin at all*, zero of 9216, comparing one plateau iteration against sixty-four. That assertion is the substance of this commit: the original test asserted a consequence (fewer basins) which a working pass need not produce, so it could have been satisfied by weakening it. A no-op check cannot pass vacuously, and it is what turned an opinion into a fact. Three candidate causes were tried and none was it. Exact float equality is genuinely wrong and is fixed regardless — a gradient computed from 8-bit samples is never exactly equal across a region the eye calls flat, so `==` never fires and `<` fires everywhere; `LEVEL_EPS` now sits behind all three comparisons. The test image is not it either: a flat disc, a terraced disc and a constant-slope ramp all behave the same. The finding worth keeping is about the domain rather than the code. On a gradient-magnitude watershed every flat region of the picture is at gradient zero, the global minimum, and a plateau with no descending exit is a minimum — one basin already, nothing to resolve. The plateaux lower-completion is defined for are regions of constant non-zero gradient, which are rarer in a photograph than F1's phrasing implies. That may be the whole answer, or it may be hiding a fourth cause; I could not close it. So `plateau_iterations` defaults to 0. The implementation stays, correct as far as it goes and costing nothing until someone finishes it; the test stays, ignored with its reason; docs/segmentation.md §12 records what was ruled out so the next attempt starts further along than this one did. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
359 lines
20 KiB
Markdown
359 lines
20 KiB
Markdown
# Region segmentation for local masking
|
||
|
||
Spec for **S15**, the spike that decides how DarkRoom finds the boundaries a local mask snaps to.
|
||
|
||
Local adjustments (FR-DEV-3, "linear gradient, radial gradient, and brush masks") need more than
|
||
placement handles to be competitive. The interactions that matter — click to select a region, drag a
|
||
contour that clings to an edge, paint without crossing a boundary — all need the same thing
|
||
underneath: **a map of where the image's regions are.**
|
||
|
||
There are two credible ways to produce that map and they are not obviously ordered. This document
|
||
specifies both, specifies the third option of combining them, and fixes the measurements that decide
|
||
between them *before* any of them is built.
|
||
|
||
---
|
||
|
||
## 1. Why this is a spike and not a build
|
||
|
||
Three properties make the choice expensive to get wrong.
|
||
|
||
**It sets the mask representation.** If regions exist, a mask is a *set of region ids* — integers,
|
||
diffable, mergeable at node level under FR-NC-9, cheap in a sidecar. If they don't, a mask is a
|
||
raster, and rasters are none of those things. This is the decision that is expensive to retrofit;
|
||
everything else in local masking sits on top of it.
|
||
|
||
**One arm collides with a settled policy.** D13 records that every dependency choice in this project
|
||
has gone the same way — rustls over aws-lc-rs, bundled SQLite, a Rust Lensfun port, zune-jpeg — to
|
||
avoid a C dependency under the Android NDK, and names `ort` as the largest exception that policy
|
||
would tolerate. Arm B needs exactly that exception. Arm A needs no new dependency at all. That
|
||
asymmetry is not a tiebreak, it is most of the cost difference, and it should be priced honestly
|
||
rather than discovered at packaging time.
|
||
|
||
**Model licensing is a distribution blocker.** D13 already establishes this for the face pipeline,
|
||
and the same reading applies here — see §7. It is a licence-reading exercise, not a research
|
||
question, and it comes first.
|
||
|
||
---
|
||
|
||
## 2. The common interface
|
||
|
||
Both arms produce the same thing. This is what makes them comparable, and what lets the choice be
|
||
deferred behind a seam rather than baked into every consumer.
|
||
|
||
```rust
|
||
/// A partition of the image into labelled regions.
|
||
pub struct RegionField {
|
||
/// Per-pixel region id at proxy resolution. R32Uint on the GPU.
|
||
labels: Texture,
|
||
/// Per-region summary: pixel count, bounding box, mean colour,
|
||
/// and (arm B only) a semantic class id.
|
||
regions: Vec<Region>,
|
||
/// Boundary strength per adjacent region pair — the edge weight
|
||
/// the merge tree is built from and the cost field reads.
|
||
adjacency: Vec<(RegionId, RegionId, f32)>,
|
||
}
|
||
```
|
||
|
||
Two consumers sit on it, and neither knows which arm produced it:
|
||
|
||
**A cost field, for contour snapping.** Live-wire — Dijkstra from the last anchor to the cursor over
|
||
a per-pixel cost that is *low* on boundaries. The cost is a **sum of terms**, which is the property
|
||
that matters: image gradient is always available, region boundary strength is added when a
|
||
`RegionField` exists, and a semantic boundary term is added when a model is present. Each source
|
||
improves the snap without changing the interface, so the arms are not exclusive here even in
|
||
principle.
|
||
|
||
**A region set, for click selection.** Click reads the label under the cursor; the mask is
|
||
`label(px) ∈ selected`. Add and subtract are set operations on ids. No flood fill, no readback, no
|
||
iteration — the whole reason the precomputed map is worth having.
|
||
|
||
Both consumers are built once, in the spike, and shared by both arms. A comparison in which each arm
|
||
gets its own consumer measures the consumers.
|
||
|
||
---
|
||
|
||
## 3. Arm A — multiscale watershed
|
||
|
||
No model, no new dependency, deterministic, works on any image.
|
||
|
||
**Gradient.** Sobel magnitude over a perceptual luma plus chroma distance, not camera-space RGB —
|
||
channel-weighted RGB gradient reads a saturated red edge as weaker than it looks. Computed after
|
||
demosaic and denoise, before the edit graph, so an exposure change does not invalidate it.
|
||
|
||
**Pre-smoothing is not optional.** Raw watershed on a noisy file makes every grain its own basin.
|
||
A guided or bilateral pre-filter, with strength tied to the file's ISO, is part of the arm rather
|
||
than a refinement of it.
|
||
|
||
**Basins.** Each pixel points downhill to its steepest neighbour; pointer-jumping resolves every
|
||
pixel to its basin root in log passes. Two compute shaders and a dispatch loop.
|
||
|
||
**The hierarchy is the cheap part.** Build the region adjacency graph, sort edges by boundary
|
||
strength, union-find over them, and *record the merge order*. That recording is the merge tree — a
|
||
click selects a leaf, and a scroll walks up through progressively coarser merges. Textbook Kruskal on
|
||
a graph of a few thousand nodes.
|
||
|
||
That node count is why the tree build is a legitimate CPU operation: it runs on the adjacency graph,
|
||
not on pixels. Pixels stay on the GPU, the graph is CPU-side — the same split ARCH §3.4 and §6.1
|
||
already draw for the edit graph, so no exception to the no-readback rule is needed.
|
||
|
||
**Known weaknesses, to be measured rather than argued about:** over-segmentation on noise and
|
||
texture, weak boundaries where contrast is low but semantics are obvious (a pale sky meeting a pale
|
||
wall), and a granularity ladder that is geometric rather than meaningful — level 7 is *a* coarser
|
||
partition, not necessarily *the* object.
|
||
|
||
---
|
||
|
||
## 4. Arm B — semantic segmentation
|
||
|
||
YOLO26-seg pretrained on ADE20K, run through `ort`.
|
||
|
||
ADE20K's 150 classes include stuff — sky, vegetation, water, wall, road — which is a far better
|
||
vocabulary for photography than COCO's 80 thing-classes. "That patch of sky" is a class here. The
|
||
nano variant is ~1.6M parameters, which is genuinely mobile-viable in a way SAM never was.
|
||
|
||
**It produces a flat partition with class ids**, so it populates `RegionField` directly: connected
|
||
components of the class map become regions, class boundaries become adjacency edges.
|
||
|
||
**What it does not produce is a hierarchy.** One partition at one semantic granularity. Click "sky"
|
||
and you get all the sky; there is no level at which you get *this part* of the sky. Adjacent
|
||
same-class regions merge whether or not you wanted them to — two different walls are one wall.
|
||
|
||
**Boundaries are semantically right and geometrically soft.** Internal stride is 4–8, upsampled to
|
||
H×W, so the class map is confident about *which* side of the boundary a pixel is on and vague about
|
||
*where* the boundary is to the pixel. Acceptable for biasing a contour. Not acceptable as a mask
|
||
edge at 100% zoom.
|
||
|
||
**Costs it brings that arm A does not:** a C dependency on the Android NDK against D13's policy, a
|
||
model to distribute and cache, an AGPL question (§7), an inference runtime per platform, and output
|
||
whose determinism across drivers is unproven (§6).
|
||
|
||
---
|
||
|
||
## 5. Arm C — semantic as a merge prior
|
||
|
||
The arms are not alternatives, and a comparison that omits their combination is a false dichotomy.
|
||
|
||
Weight each region-adjacency edge in arm A's union-find by boundary strength **and** by whether the
|
||
two regions share a semantic class. Regions that agree semantically merge earlier.
|
||
|
||
The result is a hierarchy whose coarse levels align with semantic objects and whose fine levels stay
|
||
pixel-accurate — the model doing what models are good at, which is knowing what things *are*, and
|
||
watershed doing what it is good at, which is knowing where boundaries are, exactly, at every scale.
|
||
It also repairs arm B's two weaknesses at once: the soft boundary is replaced by the watershed
|
||
boundary underneath it, and the missing granularity ladder is arm A's.
|
||
|
||
Arm C is the expected winner on quality. The question the spike actually has to answer is therefore
|
||
not "which is better" but **how much better than arm A alone, and is that increment worth D13's
|
||
cost.** §8 fixes that threshold in advance.
|
||
|
||
---
|
||
|
||
## 6. What gets measured
|
||
|
||
Per arm, over the corpus in §9, using the shared consumers from §2.
|
||
|
||
| # | Measure | Method | Why it decides anything |
|
||
|---|---|---|---|
|
||
| **M1** | **Interactions to target mask** | Clicks plus scroll steps to reach ≥95% IoU against a hand-traced mask | The real UX metric. "How many actions to get the mask I meant" is what a user experiences |
|
||
| **M2** | **Boundary accuracy** | Precision/recall of snapped-contour pixels within a 2px slack of the hand trace | Whether the edge survives 100% zoom, where masks are actually judged |
|
||
| **M3** | **Granularity coverage** | Per case, yes/no: does *any* hierarchy level produce the target region? | A hard failure mode. Arm B is expected to fail this wherever the target is not a class |
|
||
| **M4** | **Out-of-vocabulary behaviour** | M1 and M3 restricted to the OOV subset | Whether the arm degrades gracefully or produces nothing usable off-distribution |
|
||
| **M5** | **Determinism** | Same input twice on one machine; then across Mesa/AMD, NVIDIA, and Adreno | Gates whether a label field can be a cache key at all — see below |
|
||
| **M6** | **Precompute cost** | ms at proxy resolution and peak memory, on the reference desktop and one Android device | Whether it fits a background prefetch alongside the proxy |
|
||
| **M7** | **Distribution cost** | Added binary size, model size, new native dependencies, licence | D13's axis. Priced, not assumed |
|
||
|
||
**M5 deserves its own note, and it is a risk for arm A too.** ARCH §6.13 holds that cache keys are
|
||
computed over CPU-side *integer* state because GPU float results diverge across vendors. A label
|
||
field is integer state — but it is *derived from* float gradient arithmetic, so a boundary sitting
|
||
exactly between two basins could resolve differently on Adreno than on Mesa. If either arm proves
|
||
non-deterministic across vendors, its output cannot be a cache key and cannot round-trip through a
|
||
sidecar as region ids, which would push masks back toward rasters and undo most of §1's argument.
|
||
This is the measurement most likely to invalidate the whole approach, and it should be run early
|
||
rather than last.
|
||
|
||
---
|
||
|
||
## 7. Licence reading — before any code
|
||
|
||
D13's position applies unchanged: discovering at packaging time that a feature cannot ship is the
|
||
expensive failure, and it is entirely avoidable.
|
||
|
||
**Ultralytics ships YOLO under AGPL-3.0** — confirmed 2026-08-17 by reading `LICENSE` at the head of
|
||
`github.com/ultralytics/ultralytics`, which is the GNU Affero General Public License v3 verbatim.
|
||
That is deliberate on their part; the commercial licence is their business model.
|
||
|
||
GPLv3 §13 explicitly permits the combination, so this is *not* the blocker the InsightFace
|
||
non-commercial weights were: it is redistributable. But the combined work becomes effectively AGPL,
|
||
which is a change to DarkRoom's licensing posture rather than a dependency detail, and it needs to be
|
||
a decision made on purpose.
|
||
|
||
Still to verify before writing any of arm B:
|
||
|
||
- The licence on YOLO26 specifically, and on the ADE20K-pretrained weights *separately* from the
|
||
framework code — they are not necessarily the same grant.
|
||
- Whether ADE20K's own terms permit redistribution of weights derived from it.
|
||
- Whether AGPL is acceptable for DarkRoom, given Flatpak, F-Droid and Play distribution
|
||
(NFR-COMPAT-2).
|
||
|
||
**Arm A raises none of these questions**, which is worth stating plainly as part of its cost.
|
||
|
||
---
|
||
|
||
## 8. Decision criteria, fixed in advance
|
||
|
||
Stated now so the result cannot be rationalised afterwards.
|
||
|
||
- **Arm A ships alone** if it reaches within **one interaction** (M1) of arm C on the scene subset
|
||
*and* dominates arm C on the OOV subset (M4). The semantic increment does not then justify a C
|
||
dependency, an AGPL conversion, and a per-platform inference runtime.
|
||
- **Arm C ships** if it beats arm A by **two or more interactions** on the scene subset without
|
||
regressing OOV. That is a large enough difference to be felt in ordinary use, and it is what would
|
||
justify reopening D13.
|
||
- **Arm B never ships alone.** M3 is expected to fail on anything that is not an ADE20K class, and an
|
||
arm with a hard failure mode and no fallback is not a selection tool. If it surprises us and passes
|
||
M3 broadly, that is a genuine finding and this criterion is revisited on the evidence.
|
||
- **If M5 fails for an arm across vendors**, that arm cannot carry region ids into the sidecar
|
||
regardless of how it scored elsewhere.
|
||
|
||
---
|
||
|
||
## 9. Corpus
|
||
|
||
Roughly 24 images from a real library — three per category — hand-traced once and reused across all
|
||
arms. Categories chosen for the failure modes they provoke, not for coverage:
|
||
|
||
| Category | Provokes |
|
||
|---|---|
|
||
| Gradient sky | Low-contrast boundary; watershed banding |
|
||
| Foliage against sky | High-frequency boundary — arm A over-segments, arm B blurs |
|
||
| Hair against a busy background | The classic hard mask edge |
|
||
| Out-of-focus background | No edges at all; tests graceful failure in both |
|
||
| High-ISO noise | Arm A's known weakness; tests whether pre-smoothing is sufficient |
|
||
| Backlit silhouette | Strong unambiguous edge — the control case |
|
||
| Macro, abstract, still life | **OOV for ADE20K.** Arm B expected to fail M3 here |
|
||
| Architectural detail | Repeated structure; arm B merges distinct walls into one class |
|
||
|
||
Hand-tracing 24 masks is a couple of hours and it is what makes M1 and M2 mean anything. Without
|
||
ground truth this comparison is two demos and a preference.
|
||
|
||
---
|
||
|
||
## 10. Deliverables
|
||
|
||
Nothing in the UI, nothing in the graph, nothing in the sidecar.
|
||
|
||
- `core/dr-gpu/src/shaders/watershed.wgsl` — gradient and basin propagation.
|
||
- `core/dr-gpu/src/segment.rs` — the passes, producing a `RegionField`.
|
||
- RAG construction and the union-find merge tree as a pure-CPU module with unit tests and no device,
|
||
so the hierarchy is testable headless the way `dr-pipeline` is (ARCH §6.5a).
|
||
- The two shared consumers from §2 — live-wire over a summable cost field, and region-set selection.
|
||
- `core/dr-gpu/examples/segment.rs` — false-coloured PNGs at four or five hierarchy levels, plus the
|
||
M1/M2 numbers against the traced corpus.
|
||
|
||
It graduates to `core/dr-segment` if it ships; that is not a spike decision.
|
||
|
||
---
|
||
|
||
## 11. Order
|
||
|
||
1. **Licence reading (§7).** Hours, and it can eliminate arm B before anything is built.
|
||
2. **Arm A, and the shared consumers.** About a day. Look at the false-coloured PNGs — if the
|
||
granularity ladder does not feel right, nothing downstream matters and that is worth knowing
|
||
immediately.
|
||
3. **M5 across vendors, early.** It is the measurement that can invalidate the region-id
|
||
representation entirely, and it wants knowing before the corpus work is invested.
|
||
4. **The traced corpus, then M1–M4 on arm A.** Establishes the baseline every other arm is judged
|
||
against.
|
||
5. **Arms B and C**, only if §7 cleared and arm A's baseline leaves room worth closing.
|
||
|
||
Arm A is a day and needs no model, no runtime, no licence and no new dependency. It is also the
|
||
substrate every model-based arm writes into — so it is first regardless of how the comparison
|
||
eventually lands.
|
||
|
||
---
|
||
|
||
## 12. Arm A results
|
||
|
||
Built 2026-08-17. `core/dr-gpu/src/{segment.rs,hierarchy.rs}`,
|
||
`shaders/watershed.wgsl`, `examples/segment.rs`. 15 tests, 11 of them device-free.
|
||
|
||
**It works, and the hierarchy is not the expensive part.** On a 1200×800 synthetic at blur radius 2,
|
||
release build, RTX 3050 laptop: 6,730 basins and 19,223 boundaries found in **67 ms including the
|
||
readback**, and the merge tree built from them in **0.2 ms**. The tree is ~0.3% of the cost. The
|
||
estimate that priced it as a week's work was wrong by about two orders of magnitude, and the reason
|
||
is worth recording: it is Kruskal over a few thousand nodes, not a segmentation algorithm.
|
||
|
||
**The granularity ladder behaves.** At the fine end the background fragments badly — a smooth tonal
|
||
ramp bands into horizontal strips, and flat areas break into diagonal chains (see below). By
|
||
`cut_to(300)` all of that is gone: the hard-edged disc is exactly one region, the whole gradient
|
||
background is one region, and only genuine noise still fragments. The over-segmentation is absorbed
|
||
by the merge order rather than needing to be prevented, which is the property the whole design rests
|
||
on.
|
||
|
||
**Pre-smoothing is the knob it was claimed to be.** Radius 2 leaves the noisy corner fragmented at
|
||
300 regions; radius 5 largely clears it. Tying it to ISO is the right control.
|
||
|
||
Three findings that change what comes next:
|
||
|
||
**F1 — plateaux fragment into diagonal chains.** In an exactly flat region every pixel's steepest
|
||
descent is a tie, and the (value, index) tie-break sends them all up-left, so a plateau resolves into
|
||
diagonal streaks rather than one basin. Harmless here because those saddles are ~0 and the tree
|
||
merges them first — but a real sky or wall is a large plateau, and relying on the hierarchy to clean
|
||
up an artefact of the flow pass is fragile. The principled fix is a **lower-complete transform**: one
|
||
extra pass giving plateau pixels a gradient toward their nearest descending exit. Standard, cheap,
|
||
and worth doing before the corpus work.
|
||
|
||
*Attempted, and parked.* The pass exists — `plateau_init` seeds every pixel
|
||
that has a strictly lower neighbour, `plateau_step` carries a breadth-first
|
||
distance inward within a level set, and `flow` takes that distance as the
|
||
second key of a lexicographic tie-break. Bindings, ping-pong and dispatch were
|
||
all checked and are right. It is nonetheless a **measured no-op**: with a test
|
||
comparing the labelling at one iteration against sixty-four, *zero* of 9216
|
||
pixels change basin. That test is committed and ignored rather than deleted,
|
||
because it is the thing that turned "we think this works" into a fact.
|
||
|
||
Three explanations were tried and none of them was it. Exact float equality is
|
||
certainly wrong — a gradient computed from 8-bit samples is never exactly
|
||
equal across a region the eye calls flat — and a `LEVEL_EPS` tolerance now
|
||
replaces `==` and `<` in all three comparisons; it did not change the outcome.
|
||
Nor did the test image: a flat disc, a terraced disc and a constant-slope ramp
|
||
all behave identically. Worth knowing for whoever picks this up: on a
|
||
gradient-*magnitude* watershed, every flat region of the picture sits at
|
||
gradient zero, which is the global minimum, and a plateau with no descending
|
||
exit is a minimum — one basin by definition, with nothing for lower-completion
|
||
to resolve. The plateaux that do have an exit are regions of constant non-zero
|
||
gradient, which are rarer in a photograph than F1's phrasing suggests.
|
||
|
||
`plateau_iterations` therefore defaults to **0**. The pass is off, costs
|
||
nothing, and F1 stands open.
|
||
|
||
**F2 — `cut_to(N)` is a visualisation, not the interaction.** A global cut by region count spends its
|
||
budget wherever the saddles happen to be densest: at blur 5 the soft-edged disc's interior held a
|
||
cluster of near-equal saddles and ate the budget, fragmenting at a level where everything else was
|
||
clean. The real interaction walks up locally from the clicked region and has no such coupling. The
|
||
ladder in the example should not be read as what a user would experience.
|
||
|
||
**F3 — the RAG build still needs a readback.** `Segmentation::read_field` copies labels and gradient
|
||
to the CPU, gated behind the `readback` feature exactly as `read_pixels` is. Fine for a spike and
|
||
off the frame path, but a shipping build cannot take it (ARCH §6.1, AC-8), so the adjacency
|
||
accumulation has to move GPU-side with atomics. That is the largest known gap between this and
|
||
something shippable.
|
||
|
||
**M5 partially answered.** Run-to-run on one device is bit-identical, and the CPU half contributes no
|
||
nondeterminism of its own — both asserted by tests. Cross-vendor is untouched and remains the
|
||
measurement that can invalidate the region-id representation.
|
||
|
||
---
|
||
|
||
## 13. Register entries
|
||
|
||
To be added when this is accepted:
|
||
|
||
**S15** — *Region segmentation for local masking:* implement arm A and the shared consumers, trace a
|
||
24-image corpus, measure M1–M7 across arms. **Resolve the licence question before writing arm B.**
|
||
Answers: which segmentation source local masking snaps to, and whether a mask can be stored as region
|
||
ids at all. Relates to: D13, D14, FR-DEV-3, ARCH §5.4, §6.13.
|
||
|
||
**D14** — *Segmentation source for local masking* · **OPEN**. Decided by S15 against the criteria in
|
||
§8. Reopens D13's dependency-policy question if arm C wins.
|