1d7115437bd8baa7caeef392625412d7b523ef80
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d15c41e699 |
Add dr-inference-engine and route every model session through it
One crate names the runtime, the providers and the devices; dr-face and dr-segment ask it for a session by role. It hands ort an API table once per process — from a libonnxruntime it dlopens when the app names a directory holding one, otherwise from tract — so the Rust build stays free of C on every target and a package can install the runtime as a file (docs/inference.md §3). Sessions live in a registry behind a Model handle that holds the bytes, not the session: every use refreshes a timestamp and a reaper unloads whatever sat idle past the decay. A scan that runs the detector on each image never lets it go idle; a click in the develop view lets the segmenter go after thirty seconds; a handle used after that reloads, and reloads on a higher rung if a compiled engine has landed meanwhile. The probe walks the platform's ladder by building strict sessions and timing them against the CPU provider, caches the choice against a fingerprint of the runtime, driver, hardware and models, and compiles engines for the selected rung in the background, smallest model first. Nothing in this commit turns the native path on: the apps still run on tract until they call init with a runtime directory. |
||
|
|
4f4abd335f |
Cut a scene category back to the pixels that agree with it
A flag in the sky came out weighted as sky, and no feather setting fixed it. The scene model's logits are `[1, 150, 80, 80]`, so one cell is eight input pixels; at the 1600px proxy the letterbox scale is 0.4 and **one cell is 20 proxy pixels**, which `rasterise`'s bilinear then spreads across one more either side. A flag is a handful of cells whose softmax is dominated by the sky around it. The information was never in the grid, so nothing downstream of the grid can recover it. Tiling is the answer for an instance and is not available here: a category has no bounding box to tile over — sky is wherever the sky is. But the photograph is at full proxy resolution even though the weights are not, and it knows exactly where the flag is. So the model says *what*, and the pixels say *which of them*, which is the division of labour arm C already draws between the instance model and the watershed. ## Seeds, and why the erosion radius is not a guess Threshold the weights high, take `signed_distance`, and keep what is more than 1.5 cells inside. One cell *is* the model's resolution and the bilinear spreads it across one more, so the band either side of the boundary is smear rather than evidence. Deriving the radius from `Scene::cell_pixels` rather than picking a pixel count means it stays right if the proxy edge or the export changes. The mirror of that set is a confident *exterior*, free from the same field. ## Dropping small modes is the step that makes it work Four k-means modes per side, not one Gaussian: sky is blue at the zenith, white where the cloud is and pale at the horizon, and one blob over all three rejects two of them. Then modes holding under 3% of a side are discarded, and without that step the whole thing fails on the case it was built for. A small flag deep in the sky has both a high weight and a large distance from the boundary, so it lands in the interior sample and teaches the model its own colour. It cannot be excluded geometrically. It can be excluded by share. Luminance is weighted at a quarter against chrominance for the same reason the watershed's gradient is. Sky's variance is dominated by luminance, so at equal weight the distribution is a long bright streak that a mid-grey flag sits comfortably inside. A flag is separated by chrominance; a cloud is separated by luminance alone. Not zero, or a dark bird against a bright sky survives. ## Two tests, because either alone is wrong Absolute — is this colour plausible under the category, as a chi-square on the Mahalanobis distance. Comparative — is it likelier inside than outside. A pixel must pass both. The absolute test is what catches the flag, whose colour is far from *both* sides and which the comparative test alone would leave at even odds. The comparative test is what stops the absolute one needing a constant tuned per category. ## What this cannot do, written down rather than left to be discovered An intruder large enough to hold its own mode is kept. By share, a flag over a fifth of the sky and a cloud bank over a fifth of the sky are the same object, and colour does not separate them either — a white cloud is as far from blue sky in chrominance as many intruders are. So `min_cluster` is not a threshold with a correct value waiting to be found; it is the trade-off itself, set where a photographic intruder falls. Both ends are pinned by tests — `a_flag_in_the_sky_is_removed` and `an_intruder_larger_than_min_cluster_survives` — so that moving the number reads as moving the trade-off rather than as fixing a bug. The case left open is a large unrecognised object in a clean category, which wants the boundary snapped to watershed basins and is a different mechanism. ## Safe to apply without a control It is subtractive: the output is the input times a factor in `0..=1`. The worst failure available to it is losing part of a real sky, never gaining a region, so a blue car below the horizon that was never in the mask cannot be pulled into it. And a factor in `0..=1` cannot raise a sum, so `scene.rs`'s partition still holds when every category is refined independently — the weight taken off the flag lands in the unlisted remainder, which is where a flag belongs, ADE20K having no class for one. Every path without the evidence to judge returns the weights untouched and says which path it took. A refinement that silently did nothing is indistinguishable from the feature being off, and an empty seed set fitted to a distribution would reject every pixel. The signature is deliberately unchanged: categories are addressed by name, not by index, so a sharper mask cannot create the stale-index hazard the signature exists to guard against. The example writes `<prefix>-<category>-refined.ppm` beside the coarse one, never instead of it — whether this is an improvement is a comparative judgement and one image cannot answer it. Verified: fmt clean, clippy -D warnings clean, 57 dr-segment tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
df06240b2d |
Test that the category geometry means what it says
The scene tests so far checked the descriptor and the arithmetic. Neither would have noticed if `rasterise` put the sky along the bottom of the frame, because both build their own weights and never ask where those weights land. Three that do: - **Weight stays on the side it came from.** Fill the top half of the grid, read the top and bottom quarters of the image. A flipped y axis is the mistake this code is actually prone to — it is a letterbox inverse, and the numbers stay perfectly plausible when it is wrong. - **No transpose.** The vertical check alone passes under a transpose, which maps a top band onto a left band. A horizontal split is what distinguishes them, and neither test is worth much without the other. - **Coverage is a fraction.** A quarter of the cells must read 0.25. The scene tab hides a category below half a percent, so an error of a factor of the grid size would hide everything or nothing — and both look like the model failing rather than the arithmetic. `Scene::from_weights` is test-only and exists because the property under test needs weights whose correct destination is known in advance, which no real inference can provide. It uses a square window so the letterbox is the identity: any offset these find is the mapping's own rather than the padding's. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8df6000e4b |
Decode the scene model into per-category weights
The weights landed last commit with nothing to read them. This is the decoder, and the shape of it follows from one property worth stating before the code: the categories must partition the image. ## Why a partition, and not a mask per category The scene tab applies one grade to every pixel of a category — lift the sky, desaturate foliage — and both grades meet at the horizon. If each category carried an independent mask, feathering them outward would make the boundary band belong to both, so both grades would land there and every horizon would acquire a visible seam. Feathering has to *blend* there, not accumulate. So `marginalise` takes one softmax over all 150 channels and sums within each category. Grouping cannot change a total of one, so the listed categories plus the unlisted remainder sum to one at every pixel, by construction rather than by normalising afterwards. `parse_categories` refuses a descriptor that claims a class twice, because that is the one input that would quietly make the property untrue. ## The descriptor is data, and hand-written `models/scene/categories.txt` groups ADE20K's 150 classes into the eight a photographer would recognise. It is a file rather than a table in Rust for the reason `models/LICENCE.md` predicted — a vocabulary is model metadata — and it is line-oriented with comments rather than JSON like the `.classes.json` beside it, because that file is generated and this one is argued. Why `swimming pool` is water and not architecture belongs next to the line that says so. Classes are named, not indexed. An index is silently wrong after a re-export; a name is loudly wrong, and the loader refuses one the model does not have. ## Resolution, kept visible `Scene` holds the native 80×80 logit grid and resamples on demand rather than upsampling once at load. The coarseness is real — it is what the graph produces — and a type that hides it behind an early resize invites callers to expect detail that was never there. `rasterise` is where the letterbox inverse lives, once. `Letterbox` and `Window` become `pub(crate)` and `to_proto` generalises to `to_grid`, because both dense outputs this crate reads are an even fraction of the same letterboxed square and differ only in the divisor. ## Verified by looking, which is the only way this gets verified `examples/scene.rs` writes the photograph dimmed outside each category. A transposed axis or an off-by-one in the inverse produces perfectly plausible weights over slightly the wrong pixels, and no unit test catches that. On an indoor frame the person mask lands on the person, including the outstretched arm, and sky reads ~5% against a bright ceiling. It doubles as the benchmark, because every timing quoted while this model was chosen came off a laptop compiling other things and none of them belong in a document. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |