b1d1c47261936ffb657f213e7b647df33a095921
19
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
84fade99ec |
Put the developer docs under docs/dev and index the folder for users first
docs/ had 26 developer documents flat beside the manual, and the two audiences are very differently sized: most readers want the manual and the gesture reference, a few want the register, the designs and the measurements. The manual and gestures.md stay at the top; everything for someone changing the code moves to docs/dev/, and the two documents that name their own successors — the v0.1 milestone and the UI-refinement plan — go to docs/dev/archive/ rather than being deleted, since both are still cited. docs/README.md is the index, users first. Every reference follows: code comments, Cargo manifests, the workflows, the pre-commit hook, the bench and traceability tools (which locate the repo root by docs/dev/requirements.md now), packaging, the Docker READMEs, CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level deeper and is regenerated. Links out of the moved documents into the tree gain a level; a link checker over every Markdown file finds none broken. |
||
|
|
95c9cffc0d |
Keep the embedder off the Hexagon, and let the probe example ask for a runtime
On the tablet the engine compiled arcface for the NPU: the routing compared the form a rung wants with the form on offer, and for the embedder both are f32, so nothing said no. A rung now says which roles it serves at all, and the Hexagon does not serve the embedder (§7 — its vectors must compare across devices). Tested at the routing seam. dr-segment's onnx_probe example still named ort-tract, which is what stopped the workspace test build. |
||
|
|
05508741af |
Start the inference engine from both apps and show its choice in Settings
The desktop names where a package may have put libonnxruntime — an override variable, beside the executable, the package's own library directory, the Flatpak prefix, the system library directory — and Android points at the APK's native library directory, which is also what Qualcomm's DSP loader must be told for the Hexagon skel. Android starts the engine at the end of the model unpack rather than at launch, because the probe fingerprints the model files and a first launch has none until then. The About panel gains an Inference row beside Graphics, re-read every two seconds while the probe runs and engines land, and faces.model_id carries the detector's form: an int8 detector finds a different set of faces and is a different population (docs/inference.md §7). A low-memory signal drops every idle session with the GPU caches. The APK assembly bundles ONNX Runtime and the Qualcomm HTP libraries from Maven, fetched by tools/fetch-android-runtime.sh with their published checksums; RUNTIME_DIR=none builds the tract-only APK, which is a slower app and not a broken one. The desktop packages carry no runtime yet. Two probe fixes from the first desktop run: the floor must not be built with CPU fallback disabled, and a versioned libonnxruntime.so is a runtime too. On the reference desktop the probe now loads ONNX Runtime 1.30, measures 30 ms on the CPU provider, and selects TensorRT at 1.5 ms. |
||
|
|
d15c41e699 |
Add dr-inference-engine and route every model session through it
One crate names the runtime, the providers and the devices; dr-face and dr-segment ask it for a session by role. It hands ort an API table once per process — from a libonnxruntime it dlopens when the app names a directory holding one, otherwise from tract — so the Rust build stays free of C on every target and a package can install the runtime as a file (docs/inference.md §3). Sessions live in a registry behind a Model handle that holds the bytes, not the session: every use refreshes a timestamp and a reaper unloads whatever sat idle past the decay. A scan that runs the detector on each image never lets it go idle; a click in the develop view lets the segmenter go after thirty seconds; a handle used after that reloads, and reloads on a higher rung if a compiled engine has landed meanwhile. The probe walks the platform's ladder by building strict sessions and timing them against the CPU provider, caches the choice against a fingerprint of the runtime, driver, hardware and models, and compiles engines for the selected rung in the background, smallest model first. Nothing in this commit turns the native path on: the apps still run on tract until they call init with a runtime directory. |
||
|
|
42d11d919b |
cargo fmt and clippy across the panorama work, and one lint master carried
The dr-face comparison is master's: a negated partial-order test on the eye box's width, rewritten as the two conditions it meant. |
||
|
|
e4b6b6c935 |
S15.2: XFeat exports at a fixed shape and loads under tract
tools/export-xfeat.sh exports the convolutional network alone at 768×1024 grayscale, on the pattern of export-seg-model.sh: thirteen standard operator types, no dynamic axes, the keypoint decoding left to Rust. examples/onnx_probe loads it through the ort-over-tract backend the app ships with nothing unsupported and runs it in ~300 ms on the desktop CPU. The weights are Apache-2.0, read from the repository's LICENSE, with no grant on the checkpoint — recorded in models/LICENCE.md before they land, as FR-MRG-8 asks. The probe stays: the next model will need the same check. |
||
|
|
696bafa9d5 |
Undefer AI subject masking, which shipped, and give it a clause
§7 still listed "AI subject masking — deferred per D11" while MaskSource::Subject and MaskSource::Category, backed by dr-segment's instance and semantic models, had been the primary way a local adjustment is made for weeks. The code was tagged FR-DEV-3, which names gradients and brushes and says nothing about a model. FR-DEV-3i now states what exists: a subject or a category found by a local model, stored as identity with the run's signature so that it merges per field and reads as stale rather than wrong, then treated as any other layer by the edge, stroke, composition and reveal clauses. The one place it departs from FR-DEV-19 — coverage written run-length coded beside the layer, so a stored subject renders without a model — is recorded in the clause instead of left for the next audit to find. The segmentation crate and the UI's selection module are tagged to it. |
||
|
|
193b35a249 |
Start a category mask where the photograph can bear it
Clicking "architecture" made a layer whose mask was gone. Every category layer began at STRICTNESS_DEFAULT, and that constant was fitted on the synthetic sky the refine tests build — its own note warns that a real photograph's noise "moves every crossing down together", which turns out to be a considerable understatement. Measured over seven ordinary frames, half scale removes 76% to 99.5% of `architecture`, 36% to 93% of `ground` and 18% to 91% of `vegetation`. Only sky, the category the number was calibrated against, survives it. An empty mask is indistinguishable from a broken one: the layer is listed, the adjustment moves, and no pixel changes. So what this looks like from outside is that the segmentation does not make masks at all. No smaller constant fixes it either, because a nat of evidence means different things over a smooth sky and over a stone facade — the useful position is above 5 on one frame and below 1 on the next. So the frame is asked instead: `Refinement::gentle` walks down from half scale and takes the first rung whose gate removes no more than a sixth of the category's weight, and the model's own outline when none of them does. One `apply` on a friendly photograph and four on an unfriendly one, paid when a layer is made rather than for eight categories nobody masked. The slider's reset went to 4 as well, so taking the control back to its "default" emptied the mask. It goes to zero now, which is the one position documented to mean something: exactly what the model weighted. |
||
|
|
f6c9343bcc |
Ask the pixels where the edge is, not just what belongs
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m53s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 1h18m38s
Build and test / Layer separation (push) Successful in 46s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 3s
Traceability / Requirement traces (push) Failing after 50s
Build and test / Android (aarch64) (push) Failing after 30s
The colour gate decided *what* was in a category and had no way to decide *where* its edge fell. A colour test has no notion of an edge. So a refined sky lost its flag and kept the model's twenty-pixel-blocky outline, and no setting of the control could move that outline onto the horizon. This adds the second half: a **marker-based watershed**. The mask is eroded to give two markers, and the flood runs in the ribbon left between them, meeting along the most expensive line it can find. The cost is a sum of terms exactly as docs/segmentation.md §2 specifies — the photograph's own edges, and the colour model's disagreement. ## Why this is not the watershed §15 threw away That path failed because the merge *ladder* collapsed: 45,808 basins reduced to one region plus specks. There is no ladder here. Markers prevent over-segmentation by seeding rather than by merging afterwards, so the one component that broke is the one component this does not have. The markers are also better than the textbook's. scikit-image derives them by thresholding the gradient — guessing where objects are — where these come from a model that knows what sky is. Marker selection is what normally goes wrong with this method, and it was already solved. ## The gate still runs, and it runs first A flood cannot replace the colour gate. It only refines contours that already exist, and there is no contour around a flag precisely because the model never noticed one — the flag in the tests sits forty-five pixels from the boundary against a ribbon of six. The tests caught this; the first version of this commit had the flood standing in for the gate and the flag stayed. So the gate goes first and *creates* the contour, and the flood then puts every contour — the horizon and the new hole alike — onto a real edge. ## Erosion that does not delete flagpoles Eroding by a cell and a half destroys anything thinner than three cells: a mast, a bare branch, and equally a strip of sky between two of them. Those would be left unseeded and the flood would fill them from whichever side surrounds them, so a flagpole would come back — and come back *confident*. Erosion therefore stops at the ridge of the distance transform. Whatever would otherwise vanish keeps a one-pixel seed down its centre, floored at `min_thickness` so a hot pixel does not qualify. That floor also moves the signal-versus-noise decision out of colour space, where it was a share of a fitted distribution nobody can picture, and into image space, where it is a width in pixels a photographer can see. ## Two modelling errors the outward test found Both were invisible while the refinement could only subtract, because the gate was multiplied by weights that were already zero outside the mask. The moment the boundary could move outward they decided the answer. **A diagonal covariance is wrong along a gradient.** Sky moves along all three opponent features together — luminance up, red-green drifting, blue-yellow down — so treating them as independent charges a colour two deviations along that gradient three times over. Measured: sky fifteen rows past the sample scored 11.6 against a threshold of 11.34, so the model refused the very thing it was refining. The fit now carries a full 3x3 covariance, inverted by cofactors rather than by a dependency (D13, the NDK). **Eroded seeds understate the spread, always, in a known direction.** The sample is drawn from the middle of a category and never from its edge, so for anything with a gradient the colours nearest the boundary are exactly the ones left out. The broad mode is therefore fitted wider than its sample by `SHOULDER`. Same pixel: Mahalanobis 5.9 uncorrected, 1.5 corrected — the difference between refusing the horizon and reaching it. Only the broad mode is widened; the tight ones are what discriminate. ## What was given up Strict subtractivity. It bounded the damage and kept `scene.rs`'s partition true for free, and it had to go: a mask that may only shrink can sharpen a horizon inward but never outward, so wherever the coarse contour sat inside the true edge, the error survived every setting of the control. The travel bound replaces it. Everything beyond the ribbon is already a marker, so the flood never reaches it — not "can only remove" but "can only move this far", and the distance is the model's own uncertainty. That single bound also retires the connectivity test, the reachability radius and the separate additive path that an outward-growing rule would have needed. A blue car below the horizon cannot be gained, not because a rule forbids it, but because the flood is never there. `the_colour_gate_only_removes` keeps the older property where it still holds; `the_flood_cannot_travel_further_than_the_ribbon` holds the new one across the whole travel of the control. ## Cost The flood visits only unlabelled pixels, so confining it to the ribbon is not an optimisation added on top — it is what a seeded flood does. A ribbon of a few tens of pixels around one contour is a small part of a proxy. The distance transform is no longer cached, because it has to be measured from the mask as the gate leaves it and the gate moves with the control. That is one transform plus one flood per change of the control, against a precompute that runs the model once. Verified: fmt clean, clippy --workspace -D warnings clean, 63 dr-segment tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f87bf6ebc0 |
Charge a colour mode for its rarity, and keep the verdict
The refinement worked and could not be controlled. Pruning modes below a share threshold made the flag's removal a *discrete* event: below the line its Mahalanobis distance was enormous and nothing rescued it, above the line it sat at zero and nothing removed it. A control over that would appear dead through most of its travel and then start eating sky. So the prune is gone. A mode is charged `−ln(share × k)` nats, floored at zero, and that cost enters both tests — doubled in the chi-square, which is a squared distance, and directly in the log density. Rarity becomes a distance rather than a threshold, and the things a photographer wants to remove separate along it. Measured on the synthetic frame the tests build: a flag holding 1.6% of the sky is more than half gone by **2.95 nats** and a cloud bank holding a third of it survives to **5.75**. The whole interval between them is somewhere a control can sit. `the_flag_goes_before_the_cloud_does` pins the ordering, which is the property that makes one slider worth offering at all. Measured against an even split rather than against one, so raising `clusters` describes a category more finely without making every colour in it look rarer. Floored at zero so a dominant mode earns no *discount* — a bonus there would let the commonest colour outvote a bad chi-square, which is the one direction this must not bend. `Refinement` holds the per-pixel verdict, quantised to a byte over ±16 nats — an eighth of a nat per step, far finer than the narrowest transition the gate can be asked for, and the same size as the coverage buffer it sits beside. `apply` is then a smoothstep, and the model is never consulted again. That is `distance.rs`'s arrangement deliberately: there a signed distance field is computed once and feather, grow and shrink become arithmetic on it, "which is what makes those live controls rather than ones that stall on every drag". Same shape, different field. The blur moved with it, from the gate to the verdict. Smoothing the evidence rather than the decision means it is paid for once in `compute` instead of on every frame of a drag, and it is the better thing to smooth in any case. `apply` at `STRICTNESS_OFF` returns the weights untouched without reading the verdict at all. A control whose off position is *very nearly* the unrefined mask cannot answer "is this helping"; one whose off position is the unrefined mask can. `strictness_zero_changes_nothing` holds it to that, and `strictness_is_monotonic` holds the rest of the travel to only ever removing more — a slider that gave weight back partway up would be one whose direction nobody could predict. The synthetic sky is smooth enough to sit on `VARIANCE_FLOOR`, where a real one has noise and therefore a real spread, which moves every crossing down together. The ordering survives that; the placement is a calibration. Which is the honest argument for a control rather than a constant, and why the default sits at half scale instead of at the flag's measured crossing. The example sweeps the whole range and writes a frame per nat, because the question a photographer asks of a slider is where to put it, and that needs the travel rather than a point on it. Verified: fmt clean, clippy -D warnings clean, 60 dr-segment tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4f4abd335f |
Cut a scene category back to the pixels that agree with it
A flag in the sky came out weighted as sky, and no feather setting fixed it. The scene model's logits are `[1, 150, 80, 80]`, so one cell is eight input pixels; at the 1600px proxy the letterbox scale is 0.4 and **one cell is 20 proxy pixels**, which `rasterise`'s bilinear then spreads across one more either side. A flag is a handful of cells whose softmax is dominated by the sky around it. The information was never in the grid, so nothing downstream of the grid can recover it. Tiling is the answer for an instance and is not available here: a category has no bounding box to tile over — sky is wherever the sky is. But the photograph is at full proxy resolution even though the weights are not, and it knows exactly where the flag is. So the model says *what*, and the pixels say *which of them*, which is the division of labour arm C already draws between the instance model and the watershed. ## Seeds, and why the erosion radius is not a guess Threshold the weights high, take `signed_distance`, and keep what is more than 1.5 cells inside. One cell *is* the model's resolution and the bilinear spreads it across one more, so the band either side of the boundary is smear rather than evidence. Deriving the radius from `Scene::cell_pixels` rather than picking a pixel count means it stays right if the proxy edge or the export changes. The mirror of that set is a confident *exterior*, free from the same field. ## Dropping small modes is the step that makes it work Four k-means modes per side, not one Gaussian: sky is blue at the zenith, white where the cloud is and pale at the horizon, and one blob over all three rejects two of them. Then modes holding under 3% of a side are discarded, and without that step the whole thing fails on the case it was built for. A small flag deep in the sky has both a high weight and a large distance from the boundary, so it lands in the interior sample and teaches the model its own colour. It cannot be excluded geometrically. It can be excluded by share. Luminance is weighted at a quarter against chrominance for the same reason the watershed's gradient is. Sky's variance is dominated by luminance, so at equal weight the distribution is a long bright streak that a mid-grey flag sits comfortably inside. A flag is separated by chrominance; a cloud is separated by luminance alone. Not zero, or a dark bird against a bright sky survives. ## Two tests, because either alone is wrong Absolute — is this colour plausible under the category, as a chi-square on the Mahalanobis distance. Comparative — is it likelier inside than outside. A pixel must pass both. The absolute test is what catches the flag, whose colour is far from *both* sides and which the comparative test alone would leave at even odds. The comparative test is what stops the absolute one needing a constant tuned per category. ## What this cannot do, written down rather than left to be discovered An intruder large enough to hold its own mode is kept. By share, a flag over a fifth of the sky and a cloud bank over a fifth of the sky are the same object, and colour does not separate them either — a white cloud is as far from blue sky in chrominance as many intruders are. So `min_cluster` is not a threshold with a correct value waiting to be found; it is the trade-off itself, set where a photographic intruder falls. Both ends are pinned by tests — `a_flag_in_the_sky_is_removed` and `an_intruder_larger_than_min_cluster_survives` — so that moving the number reads as moving the trade-off rather than as fixing a bug. The case left open is a large unrecognised object in a clean category, which wants the boundary snapped to watershed basins and is a different mechanism. ## Safe to apply without a control It is subtractive: the output is the input times a factor in `0..=1`. The worst failure available to it is losing part of a real sky, never gaining a region, so a blue car below the horizon that was never in the mask cannot be pulled into it. And a factor in `0..=1` cannot raise a sum, so `scene.rs`'s partition still holds when every category is refined independently — the weight taken off the flag lands in the unlisted remainder, which is where a flag belongs, ADE20K having no class for one. Every path without the evidence to judge returns the weights untouched and says which path it took. A refinement that silently did nothing is indistinguishable from the feature being off, and an empty seed set fitted to a distribution would reject every pixel. The signature is deliberately unchanged: categories are addressed by name, not by index, so a sharper mask cannot create the stale-index hazard the signature exists to guard against. The example writes `<prefix>-<category>-refined.ppm` beside the coarse one, never instead of it — whether this is an improvement is a comparative judgement and one image cannot answer it. Verified: fmt clean, clippy -D warnings clean, 57 dr-segment tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
df06240b2d |
Test that the category geometry means what it says
The scene tests so far checked the descriptor and the arithmetic. Neither would have noticed if `rasterise` put the sky along the bottom of the frame, because both build their own weights and never ask where those weights land. Three that do: - **Weight stays on the side it came from.** Fill the top half of the grid, read the top and bottom quarters of the image. A flipped y axis is the mistake this code is actually prone to — it is a letterbox inverse, and the numbers stay perfectly plausible when it is wrong. - **No transpose.** The vertical check alone passes under a transpose, which maps a top band onto a left band. A horizontal split is what distinguishes them, and neither test is worth much without the other. - **Coverage is a fraction.** A quarter of the cells must read 0.25. The scene tab hides a category below half a percent, so an error of a factor of the grid size would hide everything or nothing — and both look like the model failing rather than the arithmetic. `Scene::from_weights` is test-only and exists because the property under test needs weights whose correct destination is known in advance, which no real inference can provide. It uses a square window so the letterbox is the identity: any offset these find is the mapping's own rather than the padding's. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8df6000e4b |
Decode the scene model into per-category weights
The weights landed last commit with nothing to read them. This is the decoder, and the shape of it follows from one property worth stating before the code: the categories must partition the image. ## Why a partition, and not a mask per category The scene tab applies one grade to every pixel of a category — lift the sky, desaturate foliage — and both grades meet at the horizon. If each category carried an independent mask, feathering them outward would make the boundary band belong to both, so both grades would land there and every horizon would acquire a visible seam. Feathering has to *blend* there, not accumulate. So `marginalise` takes one softmax over all 150 channels and sums within each category. Grouping cannot change a total of one, so the listed categories plus the unlisted remainder sum to one at every pixel, by construction rather than by normalising afterwards. `parse_categories` refuses a descriptor that claims a class twice, because that is the one input that would quietly make the property untrue. ## The descriptor is data, and hand-written `models/scene/categories.txt` groups ADE20K's 150 classes into the eight a photographer would recognise. It is a file rather than a table in Rust for the reason `models/LICENCE.md` predicted — a vocabulary is model metadata — and it is line-oriented with comments rather than JSON like the `.classes.json` beside it, because that file is generated and this one is argued. Why `swimming pool` is water and not architecture belongs next to the line that says so. Classes are named, not indexed. An index is silently wrong after a re-export; a name is loudly wrong, and the loader refuses one the model does not have. ## Resolution, kept visible `Scene` holds the native 80×80 logit grid and resamples on demand rather than upsampling once at load. The coarseness is real — it is what the graph produces — and a type that hides it behind an early resize invites callers to expect detail that was never there. `rasterise` is where the letterbox inverse lives, once. `Letterbox` and `Window` become `pub(crate)` and `to_proto` generalises to `to_grid`, because both dense outputs this crate reads are an even fraction of the same letterboxed square and differ only in the divisor. ## Verified by looking, which is the only way this gets verified `examples/scene.rs` writes the photograph dimmed outside each category. A transposed axis or an off-by-one in the inverse produces perfectly plausible weights over slightly the wrong pixels, and no unit test catches that. On an indoor frame the person mask lands on the person, including the outstretched arm, and sky reads ~5% against a bright ceiling. It doubles as the benchmark, because every timing quoted while this model was chosen came off a laptop compiling other things and none of them belong in a document. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e26f71d15d |
Gather every model under one tree at the repository root
The weights were in two places: face detection and recognition in `models/face/`, segmentation in `core/dr-segment/models/`. Nothing was wrong with either path, but between them there was nowhere to look to answer "how much model does this application carry", and that number is about to start growing. So the crate-local copy moves up beside the other. `models/` now holds `face/` and `segment/`, and a `du -sh` of one directory is the whole answer. No content changes: the .onnx and its vocabulary are byte-identical, and `LICENCE.md` moves up a level to cover the tree rather than one crate. The LFS pattern in `.gitattributes` is `*.onnx` and already matched both locations, so only its comment needed the new path. `include_bytes!` is relative to the source file and `build.rs` runs with the crate root as its working directory, which is why the two paths climb a different number of levels. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
56978fdf35 |
Clear the clippy warnings that were failing CI before this branch
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Desktop (Linux) (push) Failing after 9m6s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 23s
Build and test / Android (aarch64) (push) Failing after 22m38s
Nothing here is film simulation. These are lints that fail master today,
under the -D warnings CI runs with, mostly from a toolchain that learned
new ones rather than from anybody's code -- `is_multiple_of` and the
derivable `Default` did not exist as lints when this was written.
They are fixed rather than allowed, and by hand rather than by trusting
`cargo clippy --fix` wholesale: its automatic pass split a derive in two
and left a stray blank line, which is the sort of thing that is correct
and still wrong to commit.
The four that needed a decision rather than a rewrite:
- The distance transform's inner loop writes through its iterator now.
`q` stays, because it is the position the parabola is evaluated at as
well as the index it is written to -- the lint is about the write.
- `to_source` and `to_proto` take `self` by value. Their receiver is
`Copy`, so this is the same machine code and the honest signature.
- The export path's return type is five levels deep and now has a name,
plus a line saying why the `Option` wraps the `Result`: `None` is
cancellation, which is not a failure and has no error to report.
- A test fills a range instead of looping over one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
7a5e1adf51 |
Merge what two tiles saw of one subject, instead of picking a side
"Look closer" tiles the frame so a small subject reaches a fixed 640x640 model at its own size. A tile sees only the part of an object inside it, so an object on a seam produces two *partial* masks — neither of them the object. This kept the higher-scoring one and discarded the other, which quietly threw away what tiling had just been paid 2.8 seconds for: a bird with its tail cut off at a tile edge, described by whichever tile happened to hold more of the bird. Both halves existed; one survived. They are unioned now, and the overlap is what makes that sound. At 25% every pixel is seen by at least one tile at full resolution and pixels near a seam by two, so the pointwise maximum is the better estimate everywhere rather than a compromise: where one tile saw a pixel its opinion is the only one there is, and where both did, the higher value came from the tile with more context around it. A maximum of soft coverage also stays soft, which is what `prior.rs` weights merges by and what a mask layer's edge treatment needs. Each quantity gets the operation that suits it: maximum for coverage, union for the box, and the higher score rather than a blend — the score is shown to a photographer and means "how sure the model is this is a bird", so averaging in a tile that saw a wingtip would make a confident detection look doubtful for straddling a seam. The test fails against the old rule with "pixel 4 was seen by a tile and must survive the merge", which is the whole defect in one line: not a crash, not a duplicate, just a plausible mask missing half its subject. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c75849040c |
Format the tree the way the gate asks for it
`cargo fmt --check` is a required step and had drifted across 45 files. Most of it arrived this week: several operations were written in parallel worktrees and merged by hand, and a hand-merge resolves conflicts without ever running the formatter over the result. No behaviour changes — this is `cargo fmt --all` and nothing else, kept as its own commit so the next reader can skip it wholesale rather than search it for one that matters. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ec713585a5 |
Measure the distance to the edge, and get four controls for one transform
Feathering, growing, shrinking, closing and opening are the same number read differently. With the signed distance from the boundary in hand, dilation is the set where d >= -r, erosion where d >= +r, and a feather of any shape is a function of d. So the field is computed once and the controls are arithmetic on it. The **field** is what reaches the GPU, not a finished alpha, and that is the point: growing a mask or changing its falloff then costs a uniform upload and no recomputation, which is what makes them live controls rather than ones that stall on every drag. Only closing and opening rebuild, because after the first threshold the shape has changed and the old distances describe the old one. Exact Euclidean, via Felzenszwalb's separable transform — not a chamfer approximation, which leaves a mask visibly octagonal once grown more than a few pixels. A test asserts the diagonal is √2 rather than 1 or 2. It runs on the CPU, which ARCH §5.4 forbids for masks. The rule is about brush lag — a stroke rasterised per frame — and this is a different operation: once per mask edit, on input the model already produced here, producing a field the GPU then samples for free. What it buys is exact determinism, which matters because masks reach the sidecar as indices and a field that varied by vendor would mean a mask meaning one thing on the desktop and another on the phone. The half-pixel in `signed_distance` is not a detail, and a test caught it. Measuring to the nearest opposite pixel *centre* puts the smallest magnitude at 1 either side, so the boundary is nowhere and **eroding by less than a pixel removes nothing**. A control whose first notch does nothing is a broken control. Half a pixel off each side puts the boundary where it physically is, and eroding by 1 takes exactly the outermost ring. Every falloff curve is 0.5 at the boundary by construction, asserted for all five: changing the curve should change how the transition looks and never where it sits. |
||
|
|
0da8271836 |
Let the model say what a thing is and the watershed say where it ends
Local masking needs to know where an image's regions are. The watershed spike (S15 arm A) found the boundaries but had no idea what any of them enclosed; its coarse levels were geometric accidents. This adds the other half and the thing that joins them. `core/dr-segment` is where region reasoning now lives — the hierarchy moves out of `dr-gpu`, which keeps only the pixel passes that are genuinely shaders. The new crate is device-free and, without its default features, model-free too: 20 of its tests need neither an adapter nor 11 MB of weights. Arm B runs YOLO26n-seg through `ort`. D13 framed inference as a choice between `ort`'s C++ runtime and the pure-Rust dependency policy; that was a false choice. `ort`'s `alternative-backend` feature unlinks the C entirely and `ort-tract` supplies the API from tract, which is pure Rust. Measured before committing to it: zero unsupported operators, 420 ms for 640x640, and correct masks on bus.jpg. No NDK problem to solve, so D13's largest tolerated exception is not needed. Arm C is `prior.rs`, and it ships because the two arms fail in opposite directions. Instance membership re-weights the merge saddles, so region pairs the model believes share an object merge early and pairs straddling its edge merge late. No boundary moves — only the order in which they dissolve — which is how the result stays pixel-accurate at every level while its coarse levels become named things. Two things the spec assumed that turned out to be false, both recorded in models/LICENCE.md: there is no usable ADE20K-trained YOLO, so the shipped vocabulary is COCO's 80 subjects and *stuff* like sky and foliage must come from arm A; and tract cannot parse a dynamic-shape export, so the graph's input is fixed and tiling is the only route to more semantic resolution. Weights are AGPL-3.0, which GPLv3 §13 permits and which makes the combined work effectively AGPL. Deliberate, not accidental. They live in Git LFS, and a build script fails with an instruction rather than embedding a pointer file when the clone lacks them. |