Commit Graph
3 Commits
Author SHA1 Message Date
dtourolleandClaude Opus 5 f6c9343bcc Ask the pixels where the edge is, not just what belongs
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m53s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 1h18m38s
Build and test / Layer separation (push) Successful in 46s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 3s
Traceability / Requirement traces (push) Failing after 50s
Build and test / Android (aarch64) (push) Failing after 30s
The colour gate decided *what* was in a category and had no way to decide
*where* its edge fell. A colour test has no notion of an edge. So a refined
sky lost its flag and kept the model's twenty-pixel-blocky outline, and no
setting of the control could move that outline onto the horizon.

This adds the second half: a **marker-based watershed**. The mask is eroded
to give two markers, and the flood runs in the ribbon left between them,
meeting along the most expensive line it can find. The cost is a sum of terms
exactly as docs/segmentation.md §2 specifies — the photograph's own edges, and
the colour model's disagreement.

## Why this is not the watershed §15 threw away

That path failed because the merge *ladder* collapsed: 45,808 basins reduced
to one region plus specks. There is no ladder here. Markers prevent
over-segmentation by seeding rather than by merging afterwards, so the one
component that broke is the one component this does not have.

The markers are also better than the textbook's. scikit-image derives them by
thresholding the gradient — guessing where objects are — where these come from
a model that knows what sky is. Marker selection is what normally goes wrong
with this method, and it was already solved.

## The gate still runs, and it runs first

A flood cannot replace the colour gate. It only refines contours that already
exist, and there is no contour around a flag precisely because the model never
noticed one — the flag in the tests sits forty-five pixels from the boundary
against a ribbon of six. The tests caught this; the first version of this
commit had the flood standing in for the gate and the flag stayed.

So the gate goes first and *creates* the contour, and the flood then puts every
contour — the horizon and the new hole alike — onto a real edge.

## Erosion that does not delete flagpoles

Eroding by a cell and a half destroys anything thinner than three cells: a
mast, a bare branch, and equally a strip of sky between two of them. Those
would be left unseeded and the flood would fill them from whichever side
surrounds them, so a flagpole would come back — and come back *confident*.

Erosion therefore stops at the ridge of the distance transform. Whatever would
otherwise vanish keeps a one-pixel seed down its centre, floored at
`min_thickness` so a hot pixel does not qualify. That floor also moves the
signal-versus-noise decision out of colour space, where it was a share of a
fitted distribution nobody can picture, and into image space, where it is a
width in pixels a photographer can see.

## Two modelling errors the outward test found

Both were invisible while the refinement could only subtract, because the gate
was multiplied by weights that were already zero outside the mask. The moment
the boundary could move outward they decided the answer.

**A diagonal covariance is wrong along a gradient.** Sky moves along all three
opponent features together — luminance up, red-green drifting, blue-yellow
down — so treating them as independent charges a colour two deviations along
that gradient three times over. Measured: sky fifteen rows past the sample
scored 11.6 against a threshold of 11.34, so the model refused the very thing
it was refining. The fit now carries a full 3x3 covariance, inverted by
cofactors rather than by a dependency (D13, the NDK).

**Eroded seeds understate the spread, always, in a known direction.** The
sample is drawn from the middle of a category and never from its edge, so for
anything with a gradient the colours nearest the boundary are exactly the ones
left out. The broad mode is therefore fitted wider than its sample by
`SHOULDER`. Same pixel: Mahalanobis 5.9 uncorrected, 1.5 corrected — the
difference between refusing the horizon and reaching it. Only the broad mode
is widened; the tight ones are what discriminate.

## What was given up

Strict subtractivity. It bounded the damage and kept `scene.rs`'s partition
true for free, and it had to go: a mask that may only shrink can sharpen a
horizon inward but never outward, so wherever the coarse contour sat inside the
true edge, the error survived every setting of the control.

The travel bound replaces it. Everything beyond the ribbon is already a marker,
so the flood never reaches it — not "can only remove" but "can only move this
far", and the distance is the model's own uncertainty. That single bound also
retires the connectivity test, the reachability radius and the separate
additive path that an outward-growing rule would have needed. A blue car below
the horizon cannot be gained, not because a rule forbids it, but because the
flood is never there.

`the_colour_gate_only_removes` keeps the older property where it still holds;
`the_flood_cannot_travel_further_than_the_ribbon` holds the new one across the
whole travel of the control.

## Cost

The flood visits only unlabelled pixels, so confining it to the ribbon is not
an optimisation added on top — it is what a seeded flood does. A ribbon of a
few tens of pixels around one contour is a small part of a proxy.

The distance transform is no longer cached, because it has to be measured from
the mask as the gate leaves it and the gate moves with the control. That is one
transform plus one flood per change of the control, against a precompute that
runs the model once.

Verified: fmt clean, clippy --workspace -D warnings clean, 63 dr-segment tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 21:29:12 +02:00
dtourolleandClaude Opus 5 f87bf6ebc0 Charge a colour mode for its rarity, and keep the verdict
The refinement worked and could not be controlled. Pruning modes below a
share threshold made the flag's removal a *discrete* event: below the line
its Mahalanobis distance was enormous and nothing rescued it, above the line
it sat at zero and nothing removed it. A control over that would appear dead
through most of its travel and then start eating sky.

So the prune is gone. A mode is charged `−ln(share × k)` nats, floored at
zero, and that cost enters both tests — doubled in the chi-square, which is a
squared distance, and directly in the log density. Rarity becomes a distance
rather than a threshold, and the things a photographer wants to remove
separate along it.

Measured on the synthetic frame the tests build: a flag holding 1.6% of the
sky is more than half gone by **2.95 nats** and a cloud bank holding a third
of it survives to **5.75**. The whole interval between them is somewhere a
control can sit. `the_flag_goes_before_the_cloud_does` pins the ordering,
which is the property that makes one slider worth offering at all.

Measured against an even split rather than against one, so raising
`clusters` describes a category more finely without making every colour in it
look rarer. Floored at zero so a dominant mode earns no *discount* — a
bonus there would let the commonest colour outvote a bad chi-square, which is
the one direction this must not bend.

`Refinement` holds the per-pixel verdict, quantised to a byte over ±16 nats —
an eighth of a nat per step, far finer than the narrowest transition the gate
can be asked for, and the same size as the coverage buffer it sits beside.
`apply` is then a smoothstep, and the model is never consulted again.

That is `distance.rs`'s arrangement deliberately: there a signed distance
field is computed once and feather, grow and shrink become arithmetic on it,
"which is what makes those live controls rather than ones that stall on every
drag". Same shape, different field.

The blur moved with it, from the gate to the verdict. Smoothing the evidence
rather than the decision means it is paid for once in `compute` instead of on
every frame of a drag, and it is the better thing to smooth in any case.

`apply` at `STRICTNESS_OFF` returns the weights untouched without reading the
verdict at all. A control whose off position is *very nearly* the unrefined
mask cannot answer "is this helping"; one whose off position is the unrefined
mask can. `strictness_zero_changes_nothing` holds it to that, and
`strictness_is_monotonic` holds the rest of the travel to only ever removing
more — a slider that gave weight back partway up would be one whose direction
nobody could predict.

The synthetic sky is smooth enough to sit on `VARIANCE_FLOOR`, where a real
one has noise and therefore a real spread, which moves every crossing down
together. The ordering survives that; the placement is a calibration. Which
is the honest argument for a control rather than a constant, and why the
default sits at half scale instead of at the flag's measured crossing.

The example sweeps the whole range and writes a frame per nat, because the
question a photographer asks of a slider is where to put it, and that needs
the travel rather than a point on it.

Verified: fmt clean, clippy -D warnings clean, 60 dr-segment tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:40 +02:00
dtourolleandClaude Opus 5 4f4abd335f Cut a scene category back to the pixels that agree with it
A flag in the sky came out weighted as sky, and no feather setting fixed it.

The scene model's logits are `[1, 150, 80, 80]`, so one cell is eight input
pixels; at the 1600px proxy the letterbox scale is 0.4 and **one cell is 20
proxy pixels**, which `rasterise`'s bilinear then spreads across one more
either side. A flag is a handful of cells whose softmax is dominated by the
sky around it. The information was never in the grid, so nothing downstream
of the grid can recover it.

Tiling is the answer for an instance and is not available here: a category
has no bounding box to tile over — sky is wherever the sky is. But the
photograph is at full proxy resolution even though the weights are not, and
it knows exactly where the flag is. So the model says *what*, and the
pixels say *which of them*, which is the division of labour arm C already
draws between the instance model and the watershed.

## Seeds, and why the erosion radius is not a guess

Threshold the weights high, take `signed_distance`, and keep what is more
than 1.5 cells inside. One cell *is* the model's resolution and the bilinear
spreads it across one more, so the band either side of the boundary is smear
rather than evidence. Deriving the radius from `Scene::cell_pixels` rather
than picking a pixel count means it stays right if the proxy edge or the
export changes.

The mirror of that set is a confident *exterior*, free from the same field.

## Dropping small modes is the step that makes it work

Four k-means modes per side, not one Gaussian: sky is blue at the zenith,
white where the cloud is and pale at the horizon, and one blob over all
three rejects two of them.

Then modes holding under 3% of a side are discarded, and without that step
the whole thing fails on the case it was built for. A small flag deep in the
sky has both a high weight and a large distance from the boundary, so it
lands in the interior sample and teaches the model its own colour. It cannot
be excluded geometrically. It can be excluded by share.

Luminance is weighted at a quarter against chrominance for the same reason
the watershed's gradient is. Sky's variance is dominated by luminance, so at
equal weight the distribution is a long bright streak that a mid-grey flag
sits comfortably inside. A flag is separated by chrominance; a cloud is
separated by luminance alone. Not zero, or a dark bird against a bright sky
survives.

## Two tests, because either alone is wrong

Absolute — is this colour plausible under the category, as a chi-square on
the Mahalanobis distance. Comparative — is it likelier inside than outside.
A pixel must pass both.

The absolute test is what catches the flag, whose colour is far from *both*
sides and which the comparative test alone would leave at even odds. The
comparative test is what stops the absolute one needing a constant tuned per
category.

## What this cannot do, written down rather than left to be discovered

An intruder large enough to hold its own mode is kept. By share, a flag over
a fifth of the sky and a cloud bank over a fifth of the sky are the same
object, and colour does not separate them either — a white cloud is as far
from blue sky in chrominance as many intruders are.

So `min_cluster` is not a threshold with a correct value waiting to be
found; it is the trade-off itself, set where a photographic intruder falls.
Both ends are pinned by tests — `a_flag_in_the_sky_is_removed` and
`an_intruder_larger_than_min_cluster_survives` — so that moving the number
reads as moving the trade-off rather than as fixing a bug. The case left
open is a large unrecognised object in a clean category, which wants the
boundary snapped to watershed basins and is a different mechanism.

## Safe to apply without a control

It is subtractive: the output is the input times a factor in `0..=1`. The
worst failure available to it is losing part of a real sky, never gaining a
region, so a blue car below the horizon that was never in the mask cannot be
pulled into it. And a factor in `0..=1` cannot raise a sum, so `scene.rs`'s
partition still holds when every category is refined independently — the
weight taken off the flag lands in the unlisted remainder, which is where a
flag belongs, ADE20K having no class for one.

Every path without the evidence to judge returns the weights untouched and
says which path it took. A refinement that silently did nothing is
indistinguishable from the feature being off, and an empty seed set fitted
to a distribution would reject every pixel.

The signature is deliberately unchanged: categories are addressed by name,
not by index, so a sharper mask cannot create the stale-index hazard the
signature exists to guard against.

The example writes `<prefix>-<category>-refined.ppm` beside the coarse one,
never instead of it — whether this is an improvement is a comparative
judgement and one image cannot answer it.

Verified: fmt clean, clippy -D warnings clean, 57 dr-segment tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 18:30:16 +02:00