Files
DarkRoom/models/scene/categories.txt
dtourolleandClaude Opus 5 8df6000e4b Decode the scene model into per-category weights
The weights landed last commit with nothing to read them. This is the
decoder, and the shape of it follows from one property worth stating
before the code: the categories must partition the image.

## Why a partition, and not a mask per category

The scene tab applies one grade to every pixel of a category — lift the
sky, desaturate foliage — and both grades meet at the horizon. If each
category carried an independent mask, feathering them outward would make
the boundary band belong to both, so both grades would land there and
every horizon would acquire a visible seam. Feathering has to *blend*
there, not accumulate.

So `marginalise` takes one softmax over all 150 channels and sums within
each category. Grouping cannot change a total of one, so the listed
categories plus the unlisted remainder sum to one at every pixel, by
construction rather than by normalising afterwards. `parse_categories`
refuses a descriptor that claims a class twice, because that is the one
input that would quietly make the property untrue.

## The descriptor is data, and hand-written

`models/scene/categories.txt` groups ADE20K's 150 classes into the eight
a photographer would recognise. It is a file rather than a table in Rust
for the reason `models/LICENCE.md` predicted — a vocabulary is model
metadata — and it is line-oriented with comments rather than JSON like
the `.classes.json` beside it, because that file is generated and this
one is argued. Why `swimming pool` is water and not architecture belongs
next to the line that says so.

Classes are named, not indexed. An index is silently wrong after a
re-export; a name is loudly wrong, and the loader refuses one the model
does not have.

## Resolution, kept visible

`Scene` holds the native 80×80 logit grid and resamples on demand rather
than upsampling once at load. The coarseness is real — it is what the
graph produces — and a type that hides it behind an early resize invites
callers to expect detail that was never there. `rasterise` is where the
letterbox inverse lives, once.

`Letterbox` and `Window` become `pub(crate)` and `to_proto` generalises
to `to_grid`, because both dense outputs this crate reads are an even
fraction of the same letterboxed square and differ only in the divisor.

## Verified by looking, which is the only way this gets verified

`examples/scene.rs` writes the photograph dimmed outside each category. A
transposed axis or an off-by-one in the inverse produces perfectly
plausible weights over slightly the wrong pixels, and no unit test
catches that. On an indoor frame the person mask lands on the person,
including the outstretched arm, and sky reads ~5% against a bright
ceiling.

It doubles as the benchmark, because every timing quoted while this model
was chosen came off a laptop compiling other things and none of them
belong in a document.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:48:37 +02:00

56 lines
2.8 KiB
Plaintext

# Photographic categories, over ADE20K's 150 classes.
#
# The scene tab offers a slider per category, not per class: nobody wants to
# grade "sconce" and "crt screen" separately, and ADE20K's vocabulary is a
# scene-parsing benchmark rather than a photographer's list. This file is the
# translation, and it is data so that changing it is not a code change.
#
# Format: one category per line, `name = class, class, ...`, where each class
# is a name from the model's own `.classes.json`. Names rather than indices
# because an index is silently wrong after a re-export and a name is loudly
# wrong; `SceneModel::from_path` refuses a file naming a class the model does
# not have.
#
# ## Why the list is short, and why "other" is not in it
#
# Weights come from a softmax over all 150 channels summed within each
# category, so the categories listed here plus everything unlisted sum to 1 at
# every pixel. That is what lets the scene tab feather two adjacent categories
# without painting both grades into the overlap. Adding a category takes
# weight from the unlisted remainder rather than from its neighbours, so this
# list can grow without disturbing what is already here.
#
# Classes are assigned to at most one category — an overlap would break the
# partition and double-count the shared class, so the loader rejects it.
# The one the whole exercise started from. ADE20K's easiest class, and the one
# most often graded on its own in a landscape.
sky = sky
# Foliage, not "green things": a lawn and a canopy take the same saturation
# and luminance moves far more often than either takes the sky's.
vegetation = tree, grass, plant, flower, palm, field
# Standing and moving water together. `swimming pool` and `fountain` are here
# rather than under architecture because what a photographer adjusts is the
# water, not the basin.
water = water, sea, river, lake, waterfall, swimming pool, fountain
# Distant landform. Separate from `ground` because it is usually far, hazy and
# wants dehaze and contrast where a foreground surface wants neither.
terrain = mountain, rock, hill
# What the photographer is standing on, or would be. Earth and sand sit here
# rather than with terrain for the same near/far reason.
ground = earth, sand, land, dirt track, path, road, sidewalk, runway, floor, step, stairs, stairway
# Built structure. Deliberately broad: a facade, its railings and its awning
# are one surface as far as a global grade is concerned.
architecture = building, house, skyscraper, wall, tower, bridge, hovel, fence, column, awning, booth, canopy, pier, railing, grandstand
# Present for the scene tab's "expose people" move, and *not* a replacement for
# the instance model — this is every person in the frame at once, which is the
# right granularity for a global grade and the wrong one for selecting a
# subject. See `models/LICENCE.md`.
person = person