Files
DarkRoom/docs/segmentation.md
T
dtourolleandClaude Opus 5 0b20436445 Prove the plateau pass does nothing, and stop paying for it
Picks up the lower-completion work a crashed session left mid-debug, with one
failing test and no diagnosis.

The diagnosis is that the pass is a no-op. Not "does not reduce basin count" —
it changes *no pixel's basin at all*, zero of 9216, comparing one plateau
iteration against sixty-four. That assertion is the substance of this commit:
the original test asserted a consequence (fewer basins) which a working pass
need not produce, so it could have been satisfied by weakening it. A no-op
check cannot pass vacuously, and it is what turned an opinion into a fact.

Three candidate causes were tried and none was it. Exact float equality is
genuinely wrong and is fixed regardless — a gradient computed from 8-bit
samples is never exactly equal across a region the eye calls flat, so `==`
never fires and `<` fires everywhere; `LEVEL_EPS` now sits behind all three
comparisons. The test image is not it either: a flat disc, a terraced disc and
a constant-slope ramp all behave the same.

The finding worth keeping is about the domain rather than the code. On a
gradient-magnitude watershed every flat region of the picture is at gradient
zero, the global minimum, and a plateau with no descending exit is a minimum —
one basin already, nothing to resolve. The plateaux lower-completion is defined
for are regions of constant non-zero gradient, which are rarer in a photograph
than F1's phrasing implies. That may be the whole answer, or it may be hiding
a fourth cause; I could not close it.

So `plateau_iterations` defaults to 0. The implementation stays, correct as
far as it goes and costing nothing until someone finishes it; the test stays,
ignored with its reason; docs/segmentation.md §12 records what was ruled out so
the next attempt starts further along than this one did.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 12:25:28 +02:00

20 KiB
Raw Blame History

Region segmentation for local masking

Spec for S15, the spike that decides how DarkRoom finds the boundaries a local mask snaps to.

Local adjustments (FR-DEV-3, "linear gradient, radial gradient, and brush masks") need more than placement handles to be competitive. The interactions that matter — click to select a region, drag a contour that clings to an edge, paint without crossing a boundary — all need the same thing underneath: a map of where the image's regions are.

There are two credible ways to produce that map and they are not obviously ordered. This document specifies both, specifies the third option of combining them, and fixes the measurements that decide between them before any of them is built.


1. Why this is a spike and not a build

Three properties make the choice expensive to get wrong.

It sets the mask representation. If regions exist, a mask is a set of region ids — integers, diffable, mergeable at node level under FR-NC-9, cheap in a sidecar. If they don't, a mask is a raster, and rasters are none of those things. This is the decision that is expensive to retrofit; everything else in local masking sits on top of it.

One arm collides with a settled policy. D13 records that every dependency choice in this project has gone the same way — rustls over aws-lc-rs, bundled SQLite, a Rust Lensfun port, zune-jpeg — to avoid a C dependency under the Android NDK, and names ort as the largest exception that policy would tolerate. Arm B needs exactly that exception. Arm A needs no new dependency at all. That asymmetry is not a tiebreak, it is most of the cost difference, and it should be priced honestly rather than discovered at packaging time.

Model licensing is a distribution blocker. D13 already establishes this for the face pipeline, and the same reading applies here — see §7. It is a licence-reading exercise, not a research question, and it comes first.


2. The common interface

Both arms produce the same thing. This is what makes them comparable, and what lets the choice be deferred behind a seam rather than baked into every consumer.

/// A partition of the image into labelled regions.
pub struct RegionField {
    /// Per-pixel region id at proxy resolution. R32Uint on the GPU.
    labels: Texture,
    /// Per-region summary: pixel count, bounding box, mean colour,
    /// and (arm B only) a semantic class id.
    regions: Vec<Region>,
    /// Boundary strength per adjacent region pair — the edge weight
    /// the merge tree is built from and the cost field reads.
    adjacency: Vec<(RegionId, RegionId, f32)>,
}

Two consumers sit on it, and neither knows which arm produced it:

A cost field, for contour snapping. Live-wire — Dijkstra from the last anchor to the cursor over a per-pixel cost that is low on boundaries. The cost is a sum of terms, which is the property that matters: image gradient is always available, region boundary strength is added when a RegionField exists, and a semantic boundary term is added when a model is present. Each source improves the snap without changing the interface, so the arms are not exclusive here even in principle.

A region set, for click selection. Click reads the label under the cursor; the mask is label(px) ∈ selected. Add and subtract are set operations on ids. No flood fill, no readback, no iteration — the whole reason the precomputed map is worth having.

Both consumers are built once, in the spike, and shared by both arms. A comparison in which each arm gets its own consumer measures the consumers.


3. Arm A — multiscale watershed

No model, no new dependency, deterministic, works on any image.

Gradient. Sobel magnitude over a perceptual luma plus chroma distance, not camera-space RGB — channel-weighted RGB gradient reads a saturated red edge as weaker than it looks. Computed after demosaic and denoise, before the edit graph, so an exposure change does not invalidate it.

Pre-smoothing is not optional. Raw watershed on a noisy file makes every grain its own basin. A guided or bilateral pre-filter, with strength tied to the file's ISO, is part of the arm rather than a refinement of it.

Basins. Each pixel points downhill to its steepest neighbour; pointer-jumping resolves every pixel to its basin root in log passes. Two compute shaders and a dispatch loop.

The hierarchy is the cheap part. Build the region adjacency graph, sort edges by boundary strength, union-find over them, and record the merge order. That recording is the merge tree — a click selects a leaf, and a scroll walks up through progressively coarser merges. Textbook Kruskal on a graph of a few thousand nodes.

That node count is why the tree build is a legitimate CPU operation: it runs on the adjacency graph, not on pixels. Pixels stay on the GPU, the graph is CPU-side — the same split ARCH §3.4 and §6.1 already draw for the edit graph, so no exception to the no-readback rule is needed.

Known weaknesses, to be measured rather than argued about: over-segmentation on noise and texture, weak boundaries where contrast is low but semantics are obvious (a pale sky meeting a pale wall), and a granularity ladder that is geometric rather than meaningful — level 7 is a coarser partition, not necessarily the object.


4. Arm B — semantic segmentation

YOLO26-seg pretrained on ADE20K, run through ort.

ADE20K's 150 classes include stuff — sky, vegetation, water, wall, road — which is a far better vocabulary for photography than COCO's 80 thing-classes. "That patch of sky" is a class here. The nano variant is ~1.6M parameters, which is genuinely mobile-viable in a way SAM never was.

It produces a flat partition with class ids, so it populates RegionField directly: connected components of the class map become regions, class boundaries become adjacency edges.

What it does not produce is a hierarchy. One partition at one semantic granularity. Click "sky" and you get all the sky; there is no level at which you get this part of the sky. Adjacent same-class regions merge whether or not you wanted them to — two different walls are one wall.

Boundaries are semantically right and geometrically soft. Internal stride is 4–8, upsampled to H×W, so the class map is confident about which side of the boundary a pixel is on and vague about where the boundary is to the pixel. Acceptable for biasing a contour. Not acceptable as a mask edge at 100% zoom.

Costs it brings that arm A does not: a C dependency on the Android NDK against D13's policy, a model to distribute and cache, an AGPL question (§7), an inference runtime per platform, and output whose determinism across drivers is unproven (§6).


5. Arm C — semantic as a merge prior

The arms are not alternatives, and a comparison that omits their combination is a false dichotomy.

Weight each region-adjacency edge in arm A's union-find by boundary strength and by whether the two regions share a semantic class. Regions that agree semantically merge earlier.

The result is a hierarchy whose coarse levels align with semantic objects and whose fine levels stay pixel-accurate — the model doing what models are good at, which is knowing what things are, and watershed doing what it is good at, which is knowing where boundaries are, exactly, at every scale. It also repairs arm B's two weaknesses at once: the soft boundary is replaced by the watershed boundary underneath it, and the missing granularity ladder is arm A's.

Arm C is the expected winner on quality. The question the spike actually has to answer is therefore not "which is better" but how much better than arm A alone, and is that increment worth D13's cost. §8 fixes that threshold in advance.


6. What gets measured

Per arm, over the corpus in §9, using the shared consumers from §2.

# Measure Method Why it decides anything
M1 Interactions to target mask Clicks plus scroll steps to reach ≥95% IoU against a hand-traced mask The real UX metric. "How many actions to get the mask I meant" is what a user experiences
M2 Boundary accuracy Precision/recall of snapped-contour pixels within a 2px slack of the hand trace Whether the edge survives 100% zoom, where masks are actually judged
M3 Granularity coverage Per case, yes/no: does any hierarchy level produce the target region? A hard failure mode. Arm B is expected to fail this wherever the target is not a class
M4 Out-of-vocabulary behaviour M1 and M3 restricted to the OOV subset Whether the arm degrades gracefully or produces nothing usable off-distribution
M5 Determinism Same input twice on one machine; then across Mesa/AMD, NVIDIA, and Adreno Gates whether a label field can be a cache key at all — see below
M6 Precompute cost ms at proxy resolution and peak memory, on the reference desktop and one Android device Whether it fits a background prefetch alongside the proxy
M7 Distribution cost Added binary size, model size, new native dependencies, licence D13's axis. Priced, not assumed

M5 deserves its own note, and it is a risk for arm A too. ARCH §6.13 holds that cache keys are computed over CPU-side integer state because GPU float results diverge across vendors. A label field is integer state — but it is derived from float gradient arithmetic, so a boundary sitting exactly between two basins could resolve differently on Adreno than on Mesa. If either arm proves non-deterministic across vendors, its output cannot be a cache key and cannot round-trip through a sidecar as region ids, which would push masks back toward rasters and undo most of §1's argument. This is the measurement most likely to invalidate the whole approach, and it should be run early rather than last.


7. Licence reading — before any code

D13's position applies unchanged: discovering at packaging time that a feature cannot ship is the expensive failure, and it is entirely avoidable.

Ultralytics ships YOLO under AGPL-3.0 — confirmed 2026-08-17 by reading LICENSE at the head of github.com/ultralytics/ultralytics, which is the GNU Affero General Public License v3 verbatim. That is deliberate on their part; the commercial licence is their business model.

GPLv3 §13 explicitly permits the combination, so this is not the blocker the InsightFace non-commercial weights were: it is redistributable. But the combined work becomes effectively AGPL, which is a change to DarkRoom's licensing posture rather than a dependency detail, and it needs to be a decision made on purpose.

Still to verify before writing any of arm B:

  • The licence on YOLO26 specifically, and on the ADE20K-pretrained weights separately from the framework code — they are not necessarily the same grant.
  • Whether ADE20K's own terms permit redistribution of weights derived from it.
  • Whether AGPL is acceptable for DarkRoom, given Flatpak, F-Droid and Play distribution (NFR-COMPAT-2).

Arm A raises none of these questions, which is worth stating plainly as part of its cost.


8. Decision criteria, fixed in advance

Stated now so the result cannot be rationalised afterwards.

  • Arm A ships alone if it reaches within one interaction (M1) of arm C on the scene subset and dominates arm C on the OOV subset (M4). The semantic increment does not then justify a C dependency, an AGPL conversion, and a per-platform inference runtime.
  • Arm C ships if it beats arm A by two or more interactions on the scene subset without regressing OOV. That is a large enough difference to be felt in ordinary use, and it is what would justify reopening D13.
  • Arm B never ships alone. M3 is expected to fail on anything that is not an ADE20K class, and an arm with a hard failure mode and no fallback is not a selection tool. If it surprises us and passes M3 broadly, that is a genuine finding and this criterion is revisited on the evidence.
  • If M5 fails for an arm across vendors, that arm cannot carry region ids into the sidecar regardless of how it scored elsewhere.

9. Corpus

Roughly 24 images from a real library — three per category — hand-traced once and reused across all arms. Categories chosen for the failure modes they provoke, not for coverage:

Category Provokes
Gradient sky Low-contrast boundary; watershed banding
Foliage against sky High-frequency boundary — arm A over-segments, arm B blurs
Hair against a busy background The classic hard mask edge
Out-of-focus background No edges at all; tests graceful failure in both
High-ISO noise Arm A's known weakness; tests whether pre-smoothing is sufficient
Backlit silhouette Strong unambiguous edge — the control case
Macro, abstract, still life OOV for ADE20K. Arm B expected to fail M3 here
Architectural detail Repeated structure; arm B merges distinct walls into one class

Hand-tracing 24 masks is a couple of hours and it is what makes M1 and M2 mean anything. Without ground truth this comparison is two demos and a preference.


10. Deliverables

Nothing in the UI, nothing in the graph, nothing in the sidecar.

  • core/dr-gpu/src/shaders/watershed.wgsl — gradient and basin propagation.
  • core/dr-gpu/src/segment.rs — the passes, producing a RegionField.
  • RAG construction and the union-find merge tree as a pure-CPU module with unit tests and no device, so the hierarchy is testable headless the way dr-pipeline is (ARCH §6.5a).
  • The two shared consumers from §2 — live-wire over a summable cost field, and region-set selection.
  • core/dr-gpu/examples/segment.rs — false-coloured PNGs at four or five hierarchy levels, plus the M1/M2 numbers against the traced corpus.

It graduates to core/dr-segment if it ships; that is not a spike decision.


11. Order

  1. Licence reading (§7). Hours, and it can eliminate arm B before anything is built.
  2. Arm A, and the shared consumers. About a day. Look at the false-coloured PNGs — if the granularity ladder does not feel right, nothing downstream matters and that is worth knowing immediately.
  3. M5 across vendors, early. It is the measurement that can invalidate the region-id representation entirely, and it wants knowing before the corpus work is invested.
  4. The traced corpus, then M1–M4 on arm A. Establishes the baseline every other arm is judged against.
  5. Arms B and C, only if §7 cleared and arm A's baseline leaves room worth closing.

Arm A is a day and needs no model, no runtime, no licence and no new dependency. It is also the substrate every model-based arm writes into — so it is first regardless of how the comparison eventually lands.


12. Arm A results

Built 2026-08-17. core/dr-gpu/src/{segment.rs,hierarchy.rs}, shaders/watershed.wgsl, examples/segment.rs. 15 tests, 11 of them device-free.

It works, and the hierarchy is not the expensive part. On a 1200×800 synthetic at blur radius 2, release build, RTX 3050 laptop: 6,730 basins and 19,223 boundaries found in 67 ms including the readback, and the merge tree built from them in 0.2 ms. The tree is ~0.3% of the cost. The estimate that priced it as a week's work was wrong by about two orders of magnitude, and the reason is worth recording: it is Kruskal over a few thousand nodes, not a segmentation algorithm.

The granularity ladder behaves. At the fine end the background fragments badly — a smooth tonal ramp bands into horizontal strips, and flat areas break into diagonal chains (see below). By cut_to(300) all of that is gone: the hard-edged disc is exactly one region, the whole gradient background is one region, and only genuine noise still fragments. The over-segmentation is absorbed by the merge order rather than needing to be prevented, which is the property the whole design rests on.

Pre-smoothing is the knob it was claimed to be. Radius 2 leaves the noisy corner fragmented at 300 regions; radius 5 largely clears it. Tying it to ISO is the right control.

Three findings that change what comes next:

F1 — plateaux fragment into diagonal chains. In an exactly flat region every pixel's steepest descent is a tie, and the (value, index) tie-break sends them all up-left, so a plateau resolves into diagonal streaks rather than one basin. Harmless here because those saddles are ~0 and the tree merges them first — but a real sky or wall is a large plateau, and relying on the hierarchy to clean up an artefact of the flow pass is fragile. The principled fix is a lower-complete transform: one extra pass giving plateau pixels a gradient toward their nearest descending exit. Standard, cheap, and worth doing before the corpus work.

Attempted, and parked. The pass exists — plateau_init seeds every pixel that has a strictly lower neighbour, plateau_step carries a breadth-first distance inward within a level set, and flow takes that distance as the second key of a lexicographic tie-break. Bindings, ping-pong and dispatch were all checked and are right. It is nonetheless a measured no-op: with a test comparing the labelling at one iteration against sixty-four, zero of 9216 pixels change basin. That test is committed and ignored rather than deleted, because it is the thing that turned "we think this works" into a fact.

Three explanations were tried and none of them was it. Exact float equality is certainly wrong — a gradient computed from 8-bit samples is never exactly equal across a region the eye calls flat — and a LEVEL_EPS tolerance now replaces == and < in all three comparisons; it did not change the outcome. Nor did the test image: a flat disc, a terraced disc and a constant-slope ramp all behave identically. Worth knowing for whoever picks this up: on a gradient-magnitude watershed, every flat region of the picture sits at gradient zero, which is the global minimum, and a plateau with no descending exit is a minimum — one basin by definition, with nothing for lower-completion to resolve. The plateaux that do have an exit are regions of constant non-zero gradient, which are rarer in a photograph than F1's phrasing suggests.

plateau_iterations therefore defaults to 0. The pass is off, costs nothing, and F1 stands open.

F2 — cut_to(N) is a visualisation, not the interaction. A global cut by region count spends its budget wherever the saddles happen to be densest: at blur 5 the soft-edged disc's interior held a cluster of near-equal saddles and ate the budget, fragmenting at a level where everything else was clean. The real interaction walks up locally from the clicked region and has no such coupling. The ladder in the example should not be read as what a user would experience.

F3 — the RAG build still needs a readback. Segmentation::read_field copies labels and gradient to the CPU, gated behind the readback feature exactly as read_pixels is. Fine for a spike and off the frame path, but a shipping build cannot take it (ARCH §6.1, AC-8), so the adjacency accumulation has to move GPU-side with atomics. That is the largest known gap between this and something shippable.

M5 partially answered. Run-to-run on one device is bit-identical, and the CPU half contributes no nondeterminism of its own — both asserted by tests. Cross-vendor is untouched and remains the measurement that can invalidate the region-id representation.


13. Register entries

To be added when this is accepted:

S15 — Region segmentation for local masking: implement arm A and the shared consumers, trace a 24-image corpus, measure M1–M7 across arms. Resolve the licence question before writing arm B. Answers: which segmentation source local masking snaps to, and whether a mask can be stored as region ids at all. Relates to: D13, D14, FR-DEV-3, ARCH §5.4, §6.13.

D14 — Segmentation source for local masking · OPEN. Decided by S15 against the criteria in §8. Reopens D13's dependency-policy question if arm C wins.