301e6f3828d877921a1b07cfff5c4acee3c47d5f
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8df6000e4b |
Decode the scene model into per-category weights
The weights landed last commit with nothing to read them. This is the decoder, and the shape of it follows from one property worth stating before the code: the categories must partition the image. ## Why a partition, and not a mask per category The scene tab applies one grade to every pixel of a category — lift the sky, desaturate foliage — and both grades meet at the horizon. If each category carried an independent mask, feathering them outward would make the boundary band belong to both, so both grades would land there and every horizon would acquire a visible seam. Feathering has to *blend* there, not accumulate. So `marginalise` takes one softmax over all 150 channels and sums within each category. Grouping cannot change a total of one, so the listed categories plus the unlisted remainder sum to one at every pixel, by construction rather than by normalising afterwards. `parse_categories` refuses a descriptor that claims a class twice, because that is the one input that would quietly make the property untrue. ## The descriptor is data, and hand-written `models/scene/categories.txt` groups ADE20K's 150 classes into the eight a photographer would recognise. It is a file rather than a table in Rust for the reason `models/LICENCE.md` predicted — a vocabulary is model metadata — and it is line-oriented with comments rather than JSON like the `.classes.json` beside it, because that file is generated and this one is argued. Why `swimming pool` is water and not architecture belongs next to the line that says so. Classes are named, not indexed. An index is silently wrong after a re-export; a name is loudly wrong, and the loader refuses one the model does not have. ## Resolution, kept visible `Scene` holds the native 80×80 logit grid and resamples on demand rather than upsampling once at load. The coarseness is real — it is what the graph produces — and a type that hides it behind an early resize invites callers to expect detail that was never there. `rasterise` is where the letterbox inverse lives, once. `Letterbox` and `Window` become `pub(crate)` and `to_proto` generalises to `to_grid`, because both dense outputs this crate reads are an even fraction of the same letterboxed square and differ only in the divisor. ## Verified by looking, which is the only way this gets verified `examples/scene.rs` writes the photograph dimmed outside each category. A transposed axis or an off-by-one in the inverse produces perfectly plausible weights over slightly the wrong pixels, and no unit test catches that. On an indoor frame the person mask lands on the person, including the outstretched arm, and sky reads ~5% against a bright ceiling. It doubles as the benchmark, because every timing quoted while this model was chosen came off a laptop compiling other things and none of them belong in a document. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
12d320cf33 |
Record what the spec got wrong about the model that exists
docs/segmentation.md §4 priced arm B as costing a C dependency under the NDK and treated that as most of the difference between the arms. It is not a cost that has to be paid: `ort`'s `alternative-backend` disables its linking entirely and `ort-tract` supplies the API from tract, which is pure Rust. D13's "largest exception the policy would tolerate" turns out not to be needed, and the answer generalises to the face pipeline — so D13's runtime half is now answered and only its licensing half is open. Three findings contradict §4 outright and are recorded as F4-F6 rather than quietly designed around. There is no ADE20K-trained YOLO, so the shipped vocabulary selects subjects and not stuff — "select the sky" comes from the watershed or from nowhere. It is instance segmentation, so it partitions nothing and two people come back as two instances. And tract cannot parse a dynamic-shape export, which fixes the input at 640 square and makes tiling the only route to more semantic resolution. Arm C ships, but §8's criteria are not what decided it, and saying so matters more than claiming the process worked. §8 asked for a two- interaction margin over arm A on a traced corpus. That comparison was never run: F4 and F5 changed what the arms are, and a model that recognises subjects but has no word for sky cannot be a selection tool alone, while a watershed cannot tell a person from the wall behind them. They stopped being candidates and became complements. What is *not* done is written down as plainly: the 24-image corpus is untraced, so M1-M4 have no numbers and "this feels right" has not become one. M5 is answered on one device only, and region ids now reach the sidecar — so a cross-vendor divergence would mean a mask written on the desktop meaning something else on Android. F3 stands. |
||
|
|
0b20436445 |
Prove the plateau pass does nothing, and stop paying for it
Picks up the lower-completion work a crashed session left mid-debug, with one failing test and no diagnosis. The diagnosis is that the pass is a no-op. Not "does not reduce basin count" — it changes *no pixel's basin at all*, zero of 9216, comparing one plateau iteration against sixty-four. That assertion is the substance of this commit: the original test asserted a consequence (fewer basins) which a working pass need not produce, so it could have been satisfied by weakening it. A no-op check cannot pass vacuously, and it is what turned an opinion into a fact. Three candidate causes were tried and none was it. Exact float equality is genuinely wrong and is fixed regardless — a gradient computed from 8-bit samples is never exactly equal across a region the eye calls flat, so `==` never fires and `<` fires everywhere; `LEVEL_EPS` now sits behind all three comparisons. The test image is not it either: a flat disc, a terraced disc and a constant-slope ramp all behave the same. The finding worth keeping is about the domain rather than the code. On a gradient-magnitude watershed every flat region of the picture is at gradient zero, the global minimum, and a plateau with no descending exit is a minimum — one basin already, nothing to resolve. The plateaux lower-completion is defined for are regions of constant non-zero gradient, which are rarer in a photograph than F1's phrasing implies. That may be the whole answer, or it may be hiding a fourth cause; I could not close it. So `plateau_iterations` defaults to 0. The implementation stays, correct as far as it goes and costing nothing until someone finishes it; the test stays, ignored with its reason; docs/segmentation.md §12 records what was ruled out so the next attempt starts further along than this one did. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cb1d2be240 |
Choose the export folder by walking the server, not by typing it
The destination for a Nextcloud export was a text field. Nobody recalls the exact spelling of a path three levels down, and getting it wrong does not fail — `create_dir` makes whatever was typed, so a misremembered folder becomes a new one at the root and the exports are somewhere nobody looks. So it is picked the way the library root is picked, using the same `FolderBrowser` model the launch screen drives: up, into, and "use this folder", confirming the folder currently *shown* rather than one selected in the list. Same rule in both places, so the phrase means one thing. The model is shared; the worker is not. `settings_ui::spawn_folder_list` is a near-twin of the launch screen's, because that one reaches into the `LaunchController` for its session and reports onto the launch screen's error line, while this one is handed credentials and writes to the settings page. Factoring them together needs a function taking both controllers or a trait implemented twice to abstract two call sites — more machinery than the twenty lines it saves. What matters is shared already: navigation behaves identically because both drive the same model. The callbacks are wired in `lib.rs` rather than in `settings_ui::wire`, because listing a remote folder needs credentials and the settings page holds no session on purpose — it is reachable before a library is opened and must not depend on one existing. With no account the picker says to sign in first, rather than showing an empty list that reads as a server with no folders. Details that are decisions rather than accidents: the picker opens at the library root rather than at whatever half-typed path is in the field, which would list nothing and look broken. The listing area is a fixed 180px, since a folder with sixty children would otherwise push the rest of the settings page off the bottom. "Up" is disabled at the root rather than hidden, so the row does not jump as the user navigates. A failed listing leaves the picker open on the folder it was showing — where the user had got to is not something to discard over a dropped request. And the chosen folder saves immediately like every other setting on a page that has no Save button. The poll timer lives on the controller for the reason `LaunchController` keeps its own there: a `slint::Timer` stops when dropped, so one local to the function that starts it would be collected before the listing arrived. Carries in-flight work from a parallel session — a segmentation pass in dr-gpu, a sidecar cache, and the develop panel's continuing changes. 1020 tests pass, fmt clean. One clippy warning remains and is not mine: `sidecar_cache::dir` is unused while that work is in progress. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |