Commit Graph
42 Commits
Author SHA1 Message Date
dtourolle a03e082fe2 Point the manual's white balance scene at "pick", and record it again
The scene clicked 40px to the right of "pick", on "reset", so the
recording showed a neutral group being reset and a click on the wall
that panned. Re-recorded with the picker fixed: the word lights, the
sample moves temperature and tint, and Before shows what it corrected.
2026-09-20 13:41:17 +02:00
dtourolle b1c99b5796 Let the fill show the model an open void: mirror depth 0 means no ring and nothing known beyond the band 2026-09-20 10:56:38 +02:00
dtourolle d790961b28 Add the manual: every feature pictured from the application itself
Benchmarks / CPU and I/O (per commit) (push) Failing after 30s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 47s
Build and test / Layer separation (push) Successful in 27s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 4s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Traceability / Requirement traces (push) Failing after 38s
Build and test / Android (aarch64) (push) Failing after 2m24s
Build and test / Windows (x86_64, cross) (push) Failing after 3m21s
docs/manual/README.md is a tour for a photographer opening DarkRoom for
the first time — one picture per thing, moving where movement is the
point. tools/manual/ is how the pictures are made: drive.py puppeteers the
desktop build on a private Xvfb (launch, click, drag, type, screenshot,
record), scenes.py is each picture as a script, and record.sh runs them
all over a folder and writes the results into docs/manual/media/.

The media is in LFS, with the CI pulls excluding it as they exclude the
fixtures; a screenshot changes wholesale when the interface does.

Nothing in the pictures shows a person, by design: the demo library is
seventy urban and alpine frames, chosen from the catalog's rows that face
detection found nobody in.

The traceability matrix is regenerated here after the rebase that
brought this branch up to master.
2026-09-20 00:26:00 +02:00
dtourolle ecb648818b Search the user's own runtime directory before the system library
The reference desktop's only system ONNX Runtime is Arch's
onnxruntime-opt-cuda: 1.29, built without TensorRT and against cuDNN 8
on a cuDNN 9 machine. The probe rejects both providers correctly and
the app runs on the CPU provider, which is right and not what anyone
wants. runtime/ beside the models is now searched ahead of /usr/lib,
tools/fetch-desktop-runtime.sh fills it with the four libraries from
the current onnxruntime-gpu wheel (cuDNN 9, TensorRT 10), and the
About caption lists every rung that lost and why, not only the first.
Verified: the app selects TensorRT from that directory with no
environment variable set.
2026-09-19 21:15:55 +02:00
dtourolle 54a80e688c Ship MI-GAN's bare 512 generator as the panorama border filler
Sargsyan et al., ICCV 2023; MIT code and weights (models/LICENCE.md),
exported by tools/export-migan.sh at a fixed 1×4×512×512 from the
authors' checkpoint — six operator types, 28 MB, in LFS like the rest.
The package installs it beside the scene model and the APK unpacks it
with the others.
2026-09-19 20:41:20 +02:00
dtourolle 76bc5652d7 Calibrate the int8 detectors on library proxies, in chunks, and measure them
The first int8 files found no faces at all, and for two reasons the
tool now guards against. The calibration set was landscape photographs
with no faces in them, so the score head's ranges had never seen the
face regime; the set is now proxies from the library itself. And ONNX
Runtime's strided and moving-average calibration modes both degrade
these graphs measurably (a quarter of the faces at eight images, none
at ninety-six), while driving the calibrator in chunks by hand gives
ranges identical to a single pass — so the tool does that, four images
at a time, and feeds quantize_static through its range cache.

Measured against f32 over 400 proxies (docs/inference.md §10.1): the
10g form finds every face above 32 px the f32 form finds; 500m and
2.5g find 96%, and what they lose sits at a median confidence of 0.52
against the 0.50 threshold. Shipped with the number on record.

The Android unpack list gains the three int8 files; without that the
tablet never saw them. D13's runtime half records the reopening.
2026-09-19 16:02:44 +02:00
dtourolle 4ed29b9d81 Add the int8 detectors for the Hexagon, calibrated on real photographs
tools/quantise-models.sh writes the QDQ form QNN's HTP backend takes
whole: opset 17, per-channel int8 weights, uint8 activations, ranges
from running the f32 graph over photographs fed exactly as the app
feeds them. The calibration is strided, four images at a time, because
every ONNX Runtime calibrator holds each image's whole set of
activations until it folds them — a gigabyte an image on the 10g
detector, and an OOM kill with no message when folded once at the end.

Release-time, never on the device (docs/inference.md §5): it needs
real photographs and a person reading the recall measurement that
gates whether each file is offered.
2026-09-19 16:02:37 +02:00
dtourolle 05508741af Start the inference engine from both apps and show its choice in Settings
The desktop names where a package may have put libonnxruntime — an
override variable, beside the executable, the package's own library
directory, the Flatpak prefix, the system library directory — and
Android points at the APK's native library directory, which is also
what Qualcomm's DSP loader must be told for the Hexagon skel. Android
starts the engine at the end of the model unpack rather than at launch,
because the probe fingerprints the model files and a first launch has
none until then.

The About panel gains an Inference row beside Graphics, re-read every
two seconds while the probe runs and engines land, and faces.model_id
carries the detector's form: an int8 detector finds a different set of
faces and is a different population (docs/inference.md §7). A
low-memory signal drops every idle session with the GPU caches.

The APK assembly bundles ONNX Runtime and the Qualcomm HTP libraries
from Maven, fetched by tools/fetch-android-runtime.sh with their
published checksums; RUNTIME_DIR=none builds the tract-only APK, which
is a slower app and not a broken one. The desktop packages carry no
runtime yet.

Two probe fixes from the first desktop run: the floor must not be
built with CPU fallback disabled, and a versioned libonnxruntime.so is
a runtime too. On the reference desktop the probe now loads ONNX
Runtime 1.30, measures 30 ms on the CPU provider, and selects TensorRT
at 1.5 ms.
2026-09-19 16:02:37 +02:00
dtourolle 54290b9540 dr-pano: a second XFeat shape for portrait frames, and a matcher that takes seconds
Twelve real frames from the fixture set now align in 4.5 s — 4.4 s of
matching, 118 ms of bundle adjustment — where the first run took 51 s and
left the first two frames out.

The matcher computes each pair's similarity matrix once, across the
cores, with a dot product written to vectorise; both nearest-neighbour
directions read it. The frames that failed were portrait: fitted into the
landscape input they used 512 of 1024 px, and their thin overlap did not
survive at half resolution. The same weights are now exported at 768×1024
as well and the detector picks the shape by aspect. The example aligns
from embedded previews and draws the set on a cylinder; on the fixture the
sweep is 152° at a fitted 47.9 mm against the EXIF's 50, RMS 1.5 px, and
the overlaps show no ghosting.
2026-09-19 15:24:12 +02:00
dtourolle 2bf0ec8dba S15.4, CPU half: XFeat runs in ~400 ms per frame on the tablet
tools/onnx-probe-on-device.sh cross-builds dr-segment's onnx_probe
without the embedded segmentation model, pushes it with a model to the
attached device and times two runs. The 768×1024 XFeat export takes
~400 ms on the reference tablet's NEON cores against ~300 ms on the
desktop, with identical output ranges — inside NFR-MRG-1's 1 s per frame.
The blend half of S15.4 waits for a chunked blend to exist.
2026-09-19 15:24:12 +02:00
dtourolle e4b6b6c935 S15.2: XFeat exports at a fixed shape and loads under tract
tools/export-xfeat.sh exports the convolutional network alone at 768×1024
grayscale, on the pattern of export-seg-model.sh: thirteen standard
operator types, no dynamic axes, the keypoint decoding left to Rust.
examples/onnx_probe loads it through the ort-over-tract backend the app
ships with nothing unsupported and runs it in ~300 ms on the desktop CPU.

The weights are Apache-2.0, read from the repository's LICENSE, with no
grant on the checkpoint — recorded in models/LICENCE.md before they land,
as FR-MRG-8 asks. The probe stays: the next model will need the same
check.
2026-09-19 15:24:10 +02:00
dtourolle 6aae4c3eb0 Ship the three eye-state models beside the face pair
2d106det for the eye contours, OCEC for open or closed, SGC for
sunglasses — all three pinned to a batch of one by the same script as
the pair, and installed by every packager so the eyes-open filter works
out of the box. The two classifiers are MIT, code and weights; the
README records their provenance, SGC's undocumented training set, and
the hashes as fetched and as shipped.
2026-09-19 14:04:08 +02:00
dtourolle c826fed605 Put the plugin API post-v1, and let the matrix count it that way
The register said two things about plugins. §7 had listed "Plugin API"
as deferred since the first draft, in a bare row; §3.10 then specified
it in 23 clauses that counted against coverage. Twenty-one of them had
no implementation of any kind, and could not have: no crate loads
anything at runtime. The coverage figure was measuring the contradiction.

Decided 2026-09-19: §7 is right. §3.10 stays as the design of record,
each of its clauses is marked "(post-v1)" on its defining line, and
NFR-SEC-6 — which exists only for plugins — goes with them, as does D16.

The traceability tool learns the marker. A deferred requirement is still
defined, so a tag naming it is not an orphan, but it leaves the
denominator and is listed in its own table rather than under "not yet
tagged". The marker must sit on the definition line; a mention of
"post-v1" in prose changes nothing, and where an ID is defined twice the
deferral on either line wins. Both are tested. Coverage moves from 72.2%
of 194 to 80.6% of 170 without a line of application code changing,
which is the honest figure: it now measures what v1 owes.
2026-09-19 12:25:03 +02:00
dtourolle 4f31123b0c Let the user choose which SCRFD finds their faces
faces.md §12.3 measured what the cheapest detector costs: the small
faces in every group shot, and a dog embedded a dozen times. Which
trade is right depends on the machine doing the sweep — a desktop left
overnight and a tablet on a battery want different answers — so the
detector is now a per-device setting, Fast / Balanced / Thorough on
the settings page beside the indexing button, persisted with the rest
of the settings file.

A detector is half of a model id. Every face, marker, shard and
calibration is keyed on faces.model_id precisely so that a model change
is a new id and a re-index rather than a silent change under existing
data, and a detector change is a model change: it decides which faces
exist and where the landmarks that align them land. So each choice
names its own pipeline. 500M keeps the bare "w600k_mbf" every existing
library was written under, so an upgrade disturbs nothing; the others
are qualified. Choosing one restarts coverage from zero under the new
id, the sweep re-detects, confirmed names carry across by box overlap,
and the sync shards are keyed by the same id so a peer on another
setting neither adopts nor pollutes them. The library controller
carries the id into the sync the same way it carries the cache budget,
because the sync starts from places that have no settings in reach.

All three shape-fixed exports ship — APK, Arch, Flatpak — since a
tablet has no other way to obtain the one it was not installed with;
the APK grows by twenty megabytes for the choice.
2026-09-11 22:12:53 +02:00
dtourolle d8b9b5a4bb Let the release script write the release commit it was never trusted with
Every release commit in the history reads `Release X.Y.Z` and none of them
carries the message this script would have written, so nobody has ever
passed it `--commit` — and the reason is in the message it wrote: a
Co-Authored-By trailer naming an assistant, which no commit in this repository
carries and none should.

The trailer goes, and so does the paragraph above it: the script's own header
already says why it exists, and a release commit is the one place a one-line
subject is the whole convention.
2026-09-11 09:33:28 +02:00
dtourolle c48896bd95 Merge: touch selection and drag, from the gallery-selection branch
Verified before merge: fmt clean, clippy -D warnings clean, 563 dr-ui tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	docs/traceability.md
2026-08-30 13:45:18 +02:00
dtourolleandClaude Opus 5 745f3a98c7 Merge: measure the catalog, so §8's promise stops being a promise
There was no benchmark harness of any kind -- no benches/, no criterion,
no synthetic fixture -- while §8 promised a suite run per commit that
fails the build on regression. Ten performance requirements could be
neither passed nor failed.

tools/bench builds a deterministic 50,000-row catalog over a pool of
twelve generated JPEGs, about 14 MB, reproducible from a seed, with a
stamp so it rebuilds rather than silently comparing against a different
workload. It depends on nothing GPU or UI, which is what makes the CI job
affordable.

NFR-P1 and NFR-P3 are gated and tagged. NFR-P7, NFR-P8 and R2 are
measured but deliberately untagged: the export gate is one-sided, the
memory figure is the catalog layer's share rather than the whole, and
R2's first sentence is a 60 fps scroll a catalog benchmark cannot claim.

Every recorded value in the baseline is null. Nobody has run this on the
reference desktop, and a fabricated figure would make every later
comparison a comparison against a guess.

First run on this machine: catalog opens in 70 ms against a 2 s budget,
and thumbnail throughput measures 37 img/s against a target of 100 --
reported rather than asserted here, and the first evidence that the
target may not hold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 13:13:57 +02:00
dtourolleandClaude Opus 5 a8008feaf3 Inline the verdict header, and take the formatter's pass
clippy::print_literal on the results table header, and rustfmt's first
look at code whose author could not run cargo. The four constructs the
author flagged as risky -- scoped-thread lanes, a seventeen-argument
params!, is_some_and over a closure, a refutable let-else -- all compiled
untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 13:13:57 +02:00
dtourolle e8e96eed40 Measure the performance targets §8 has been promising, and fail on a regression
docs/requirements.md §8 has said since it was written that performance is
verified by "an automated benchmark suite against a synthetic 50k catalog, run
per-commit … A regression beyond stated tolerance fails the build." There was
none. No benches/, no [[bench]], no criterion, no synthetic catalog, and three
CI workflows that between them measured nothing. Ten performance requirements
could therefore be neither passed nor failed, and five of them carried a
TRACES: tag regardless.

tools/bench is the half of that promise that can be kept honestly on a runner
with no GPU and no display.

# The fixture

Rows are cheap and pixels are not, so it builds fifty thousand catalog rows
over a pool of a dozen real files, each referenced by several thousand of them.
Everything the catalog half touches is rows and is exact at full scale;
everything the pixel half touches is one file at a time and does not care how
many rows point at it. Fourteen megabytes on disk instead of two terabytes, and
neither half is flattered by the trade. It is reproducible from a seed, and a
stamp beside it — seed, row count, source size, dr-catalog's schema version —
rebuilds it rather than letting a run be compared against a baseline that
describes a different library.

# What it can now pass or fail

NFR-P1, and R2's second sentence with it: Catalog::open plus the count, first
window and timeline the grid cannot paint without. The interesting part turned
out to be the open itself — schema::backfill runs three passes over the images
table on every open, which is O(library) work on a path whose budget is stated
in absolute seconds. Tagged TRACES: NFR-P1, on a gate that fails if it breaks.

NFR-P3: thumbnail throughput on the embedded preview path, through the same
per-image work spawn_thumbnail_sweep does and in the same shape — chunks of 96,
lanes owning disjoint slices, the single thread that owns the store writing the
finished chunk. Mirrored rather than called, because that function takes a
RemoteBackend and would measure somebody's network. Tagged TRACES: NFR-P3.

# What it deliberately does not claim

NFR-P7 is the whole chain, and only the encode half of it runs without an
adapter. So the export row is a one-sided gate — over two seconds in the encode
alone violates the requirement; under it proves nothing — and there is no
TRACES: NFR-P7 anywhere. NFR-P8 is about the application at idle, and the probe
is a process holding the catalog and nothing else, so it records the catalog
layer's share and carries no budget until somebody decides what that share
should be. No tag there either. CONTRIBUTING.md asks that a requirement be
closed by a test that would fail if the behaviour were removed, and two more
plumbing tags is what this repository already has too many of.

NFR-P8 also gets the answer §4.1 demands: RSS is exclusive of device-local GPU
allocations and cannot be made otherwise, because such an allocation never
enters the process's address space. The requirement should be restated as two
figures, and docs/benchmarks.md says so.

# Two gates, and why one of them steps aside off the reference desktop

The budget is the requirement's own number and never moves. The baseline is
what the reference desktop last measured, and drifting 15% past it fails the
build even while still inside the budget — which is how performance rot
actually arrives, never over the line, always a little worse.

A budget written for twenty-four threads cannot be asserted on a two-core
container. §8 names the reference desktop, not CI, so each metric declares
whether its budget is machine-sensitive; those are asserted under --reference
and reported everywhere else. Catalog open is not one of them: two seconds
against an expected figure two orders of magnitude smaller is a threshold any
machine can be held to. This is the trap core/dr-gpu/tests/frame_budget.rs
already refuses — a red gate everybody learns to ignore.

# The baseline ships with no numbers in it

Every recorded field is null, because nobody has run it yet. Writing
plausible-looking figures would make every later comparison a comparison
against a guess, and the first real regression would be invisible. Run
`dr-bench record --reference` on the reference desktop and commit the diff;
until then the budget gate works and the report says the other one cannot.

# CI

.gitea/workflows/benchmark.yml, and its own workflow rather than a step in
build-and-test.yml: a red "Build and test" says the code is wrong, a red
"Benchmarks" says it got slower, and the second must not be reachable by
retrying a flaky compile. The cpu job runs on every push and builds -p dr-bench
alone — which is why that crate depends on no GPU and no UI crate. The gpu job
is the frame budget that already exists and already skips without an adapter,
on workflow_dispatch, because building wgpu on every commit to rediscover that
the runner has no device is not a use of anybody's minutes.
2026-08-30 10:40:10 +02:00
dtourolleandClaude Opus 5 6ec6c5cbd6 Count the tags this rule would have cost, rather than guessing at it
The module doc said an extractor keyed on string literals would "lose six
genuine tags to save two false ones". The six is right — `schema.rs` carries
that many `-- TRACES:` lines inside Rust literals — but the two undercounted
the false ones, which were R1 twice, FR-CAT-1 twice, FR-CAT-2, NFR-P1, one in
`gestures.rs` and two emitted by `dr-pipeline/build.rs`. The comparison it was
drawing does not need a number on that side to hold.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:25:13 +02:00
dtourolleandClaude Opus 5 0906c3983f Stop reading a coverage calculation as a diagnostics subsystem
NFR-OPS-1 asks for structured levelled logging to a rotating, size-capped
on-disk log in the XDG state or Android app directory, automatic redaction of
credentials and tokens, and a one-click diagnostics bundle with an explicit
preview-and-consent step. Its two tags were on `compute_coverage` and on the
traceability tool's gesture extractor.

Neither is diagnostics under any reading. One computes a ratio and the other
generates a markdown document; neither writes a log, and no rotating on-disk
log exists anywhere in the tree — logging goes to stderr and to logcat.

These were real tags, not the fixtures the extractor was just taught to
ignore, which makes them the more instructive case: the tool was correct and
the tags were wrong. NFR-OPS-1 is untagged again, and outstanding.md §9 now
says what is actually missing rather than that the requirement is covered.

`gestures.rs` keeps its FR-UI-4 tag, which is a separate claim and unaffected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:17:53 +02:00
dtourolleandClaude Opus 5 9d9abb227f Ask where a TRACES tag sits, so the tool stops tagging its own fixtures
The traceability tool scans `tools/`, which is its own source, and a line was
taken for a tag whenever `TRACES:` appeared anywhere on it. Its unit-test
fixtures are therefore tags. R1 — cross-platform output within a bounded
tolerance, the requirement with no acceptance criterion at all — was reported
implemented on the strength of two string literals in
`context_looks_forward_then_backward`.

R1 was only the visible case because it had no other coverage. The same
fixtures also contributed sites to FR-CAT-1, FR-CAT-2 and NFR-P1, `gestures.rs`
contributed one to FR-UI-4 from a `push_str`, and `dr-pipeline/build.rs`
contributed FR-DEV-3a and FR-DEV-3c from the tag it *emits* into generated
code. Those four requirements keep real tags elsewhere, so nothing but noise is
lost by dropping them.

The rule is about position, not about string literals. It cannot be about
string literals: `schema.rs` writes six genuine tags inside Rust string
literals, because the SQL it embeds is commented with `--`, and an extractor
that refused those would lose more than it saved. What separates the two is
where on the line the tag is. A tag written to be read is the first word of its
comment; a tag quoted inside an expression never is. So `tag_body` asks for a
comment opener at the start of the line and `TRACES:` immediately after it.

That closes every shape but one: a multi-line literal whose lines really do
begin with `///`, which no line-oriented reader can tell from source. There is
one such fixture and its ids are now UT and IT, which `is_requirement` already
excludes from coverage — the mechanism existed and was simply never used on
the tool itself. `this_crates_own_fixtures_cannot_reach_the_register` enforces
that: any requirement id below `mod tests` in this crate fails the test and
says to use a UT- or IT- id instead. Tagging the tool's real code is still
allowed.

On `SOURCE_SUFFIXES`, which cannot reach `AndroidManifest.xml`, the Flatpak
manifest, the Dockerfile or the CI workflows: it is deliberately left alone,
and the reasoning is recorded beside it. A tag on a manifest asserts that a
comment exists next to a line nothing checks, which is the weak form
CONTRIBUTING.md warns about. The convention already in the tree — a Rust test
that `include_str!`s the file and asserts what must be in it, with the tag on
the test — is what a tag is supposed to mean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:17:34 +02:00
dtourolleandClaude Opus 5 9a2b39b8e5 Add the ADE20K scene model beside the instance one
`models/LICENCE.md` recorded, on 2026-08-21, that no YOLO model trained on
ADE20K existed in usable form — the stuff classes photography cares about,
sky and vegetation and water, had no model to come from. Re-checked
2026-08-30: Ultralytics now ships a `semantic` task with ADE20K
checkpoints, so `models/scene/` holds `yolo26s-sem-ade20k`.

This is an addition, not a replacement. A semantic model labels every
pixel but merges same-class pixels into one region, so it cannot tell
three people apart — which is exactly what clicking a subject needs, and
exactly what `segment/`'s COCO instance model already does. The scene tab
grades per category and does not care that instances are merged. Keeping
both is the point.

## The export is truncated, deliberately

Ultralytics ends the graph with `Resize -> ArgMax -> Cast` and hands back
a `[1, 640, 640]` u8 label map. The script cuts that tail and exposes the
classifier's `[1, 150, 80, 80]` f32 logits instead, for two reasons.

Cost: the Resize materialises 150 x 640 x 640 x f32, 246 MB, and ArgMax
then reduces across the channel axis, striding 409,600 elements per
comparison. On one loaded machine the full graph ran ~1160 ms against
~500 ms truncated — roughly four fifths of the time spent on work the
application discards. Those numbers were measured under contention and
are upper bounds, but the ratio is structural.

Softness: ArgMax destroys the per-class scores, and the scene tab needs
them. Softmax over the 150 channels, summed within each photographic
category, yields per-category weights summing to 1 at every pixel.
Feathering a partition of unity cannot double-grade a boundary, whereas
feathering hard labels outward from two adjacent categories paints both
grades into the overlap and haloes every horizon.

The discarded upsample was never information: the graph's true spatial
resolution is the 80x80 logit grid, and the application can resample from
that itself.

The tail is matched by op type and asserted before cutting, so an
upstream graph change fails loudly in the exporter rather than quietly
shipping a differently-shaped model.

Nothing reads these weights yet — the decode path, the category
descriptor grouping 150 classes into ~8 photographic ones, and the scene
tab are still to come. At 24 MB this model also wants the runtime-asset
treatment `models/face/` already gets on Android rather than
`include_bytes!`; embedding it would put ~35 MB of weights in the binary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:05:44 +02:00
dtourolleandClaude Opus 5 2e09906a08 Keep the export venv off tmpfs
`mktemp -d` lands in `/tmp`, which on this and most current Linux
distributions is a tmpfs — memory, not disk, sized at half of RAM. The
venv this script builds installs torch into it, several gigabytes, and
the failure mode is not subtle:

    error: Failed to install: torch-2.13.0-...whl
      Caused by: No space left on device (os error 28)

on a machine with 102 GB free on the filesystem holding `/var/tmp`. The
quieter version of the same bug is worse: when it does fit, it evicts
whatever the user had in page cache to make room.

`${TMPDIR:-/var/tmp}` respects an explicit TMPDIR and otherwise picks the
disk-backed directory, which is what a multi-gigabyte throwaway wants.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:05:19 +02:00
dtourolleandClaude Opus 5 e26f71d15d Gather every model under one tree at the repository root
The weights were in two places: face detection and recognition in
`models/face/`, segmentation in `core/dr-segment/models/`. Nothing was
wrong with either path, but between them there was nowhere to look to
answer "how much model does this application carry", and that number is
about to start growing.

So the crate-local copy moves up beside the other. `models/` now holds
`face/` and `segment/`, and a `du -sh` of one directory is the whole
answer.

No content changes: the .onnx and its vocabulary are byte-identical, and
`LICENCE.md` moves up a level to cover the tree rather than one crate.
The LFS pattern in `.gitattributes` is `*.onnx` and already matched both
locations, so only its comment needed the new path.

`include_bytes!` is relative to the source file and `build.rs` runs with
the crate root as its working directory, which is why the two paths climb
a different number of levels.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 10:05:04 +02:00
dtourolleandClaude Opus 5 091306f736 Gate the gesture vocabulary the way the matrix is gated
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Desktop (Linux) (push) Failing after 1h16m50s
Build and test / Layer separation (push) Successful in 39s
Traceability / Requirement traces (push) Successful in 50s
Build and test / Android (aarch64) (push) Successful in 22m27s
Three holes, all found by the gate catching itself out.

**Prose that mentions the tag was read as a tag.** `dr-ui`'s module list
carries a comment saying where the generated table comes from, and it
names `GESTURE:` in passing; the scan extracted that sentence fragment as
a gesture with no place and no way to perform it. A tag must now *open*
its comment. A line that merely mentions it is describing the mechanism,
not declaring a member of it, and position is the only thing that tells
the two apart — which also makes the string-literal guard fall out for
free rather than being a special case.

**Neither artefact was regenerated on commit.** They cite line numbers,
so they go stale on anything that moves a line — the sheet commit made
the document wrong about every gesture in `library.slint` without
touching a single one. The pre-commit hook that already keeps the matrix
in step now keeps these too, and unlike the matrix it *fails* rather than
shrugging when the scan does: a matrix that will not build leaves a stale
one in place, where a malformed gesture block means a user about to be
told the wrong thing.

**CI did not check them at all.** It does now, blocking. The matrix is
read; the gesture table is *shown to somebody using the application*, and
a stale one tells them to perform a gesture that no longer exists — from
which they will conclude the application is broken rather than the page.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 00:51:13 +02:00
dtourolleandClaude Opus 5 31a3580f9d Extract the gesture vocabulary from the code that implements it
Every gesture the application has was documented in the comment beside
the `TouchArea` that implements it. Excellent comments, and unreachable
by anyone not reading the source — which is the FR-UI-4 failure in a
different costume: a gesture nobody can find is a feature only its author
knows about.

Writing them out again in a hand-kept help page is the failure this
avoids. Two descriptions of one gesture drift, and it is always the prose
that drifts: the code is exercised every time somebody uses the
application and the page is exercised never. A help screen confidently
describing a double tap the grid stopped honouring last week is worse
than no help screen — and the grid did stop honouring one, in the commit
before this.

So the comment beside the implementation stays the only copy, and a
`GESTURE:` block beside it is scanned into two artefacts: `docs/gestures.md`
for a reader, and a Rust table for the application to draw a help sheet
from. Both committed, both gated, so neither can quietly stop describing
the code.

It lives in the traceability crate because it is the same operation on
the same input — walk the tree, pull structured tags out of comments,
render, fail if the committed artefact has moved. Only the vocabulary is
new. It scans `ui` and `apps` alone: a gesture needs an interface to be
performed on, and excluding `tools` is also what stops the scanner
extracting its own worked examples as broken gestures.

Fifteen gestures so far, across the library grid and the People screen.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 00:09:49 +02:00
dtourolleandClaude Opus 5 f4395bd17c Run the face tests on the tablet, where the NEON kernel actually runs
The similarity scan picks its dot product per machine, and the NEON one
is the kernel that ships to the phone and the tablet — and the one a
desktop cargo test never executes. A wrong lane index or a mishandled
tail there is a silent wrong answer on exactly the devices nobody runs
the suite on, which is a poor place for the only untested code path.

dr-face carries no weights and touches no display, so its tests are a
plain ARM64 binary that runs under adb shell with nothing installed.
The script builds it against the SDK's newest NDK, pushes it, runs it
and cleans up. It checks for the device first, so a tablet that is not
plugged in costs a second rather than the two minutes it takes to
compile for it.

Not wired into CI, which has no device attached.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 12:02:03 +02:00
dtourolleandClaude Opus 5 69dccaa061 Count the platform layer's traceability tags
`platform` was missing from the traceability scanner's source roots, so
every TRACES tag in `dr-plat` — the secret store, volume discovery, and
now display-profile acquisition — was invisible to the matrix. The
FR-PLAT-* family is exactly what that crate exists to satisfy, so the
omission understated coverage by the requirements it was meant to
count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 18:46:59 +02:00
dtourolleandClaude Opus 5 72410f39c6 Answer M1: tract loads both face graphs once their dims are pinned
Neither InsightFace export parses as shipped -- SCRFD fails at its input
node, ArcFace at the first Conv -- which is the same wall dr-segment hit
on YOLO's dynamic export. Both load cleanly with the input dims frozen,
so the pure-Rust runtime holds for the face pipeline too.

tools/fix-face-model-shapes.sh does the freezing, and exists so the
artefact is reproducible rather than a binary someone once produced. It
takes two forms because the two graphs need different ones: ArcFace's
batch is a named dim_param, SCRFD's H and W are dynamic but unnamed.

Also notes YuNet loading with no intervention, which matters for the
licence question in faces.md 2.3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 19:51:02 +02:00
dtourolleandClaude Opus 5 ff1a3e0e80 Print the black-and-white negatives instead of showing the scan
Choosing Ilford HP5 Plus showed an inverted grey frame. So did Double-X.
They are negatives, and upstream leaves `target_print` null on every
monochrome stock, so nothing was ever printed and the scan was all there was.

A colour negative at least announces itself -- the orange mask says plainly
that you are looking at a negative. A monochrome one just looks broken.

They print on Kodak 2302 now, which is a monochrome print film and is what
such a negative is actually printed onto; Double-X onto 2302 is the standard
cine chain. For the Ilford stocks it stands in for an Ilford paper, which
nobody has measured, and is at least the right kind of material.

The scan is still reachable through the Scanned/Printed toggle. It is a thing
to choose now rather than the only thing on offer.

`every_shipped_stock_bakes` did not catch this, and could not: it derives
"should this be inverted?" from the stock's kind *and whether it names a
paper*, so it looked at an inverted HP5, concluded that was right for an
unprinted negative, and passed. The assertion was self-consistent and the
situation was still wrong. The new test asserts the thing that actually
matters -- a camera negative must name a paper, that paper must be a printing
stock, and it must be the same kind of material, so a monochrome negative
cannot end up on colour paper.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 17:35:25 +02:00
dtourolleandClaude Opus 5 ce6458547a Develop longer, from the measurements rather than from a contrast slider
Pushing was not a thing to simulate. It was measured data being thrown
away: Double-X and 2302 each ship five characteristic curves, one per
development time, and this shipped the 6.5-minute column and discarded
four. All five now ship and interpolate.

The axis is real. Double-X runs 4 to 12 minutes, and across it the average
gradient goes 0.472 to 1.034 while Dmax goes 1.19 to 2.56.

The control is in stops, because that is what a photographer means, and one
stop is a factor of about 1.41 in time. That mapping is checked rather than
assumed: against Double-X's own axis it lands within 2% of the 9-minute
column for +1, and near 12 minutes for +2, which are the times the datasheet
gives for exactly that. There is a test.

**Pushing must not recover shadow detail, and this does not.** Across the
whole measured range the speed point moves about a third of a stop while the
gradient doubles; three stops under mid-grey, density goes from 0.008 to
0.035, which is still nothing. Developing longer multiplies what was already
recorded and cannot record what never hit the film. A push built as added
exposure or global contrast brightens those shadows instead and looks
convincing until someone who shoots film sees it, so that property has a
test of its own.

Interpolated in *log* time, because development is multiplicative: 4 to 5
minutes is the same amount of push as 9 to 12, and interpolating linearly
would bunch the control at one end. Clamped at both ends, because past the
published range there is no data and extrapolating a contrast curve invents
an emulsion nobody tested. A stock measured at one process ignores the
control entirely rather than inventing a curve for it -- Portra 800's pushes
are separate *measured* profiles, which is the honest way to offer those.

Costs nothing per pixel and changes no shader. The curves are a per-stock
table, so the interpolation happens on the CPU at bake time, where choosing a
stock and moving its sliders already rebakes. The Vulkan shader is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 17:35:25 +02:00
dtourolleandClaude Opus 5 e9b3598841 Add five Ilford stocks, and say plainly that they are constructed
Build and test / Desktop (Linux) (push) Successful in 20m48s
Build and test / Layer separation (push) Successful in 27s
Traceability / Requirement traces (push) Failing after 25s
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 2s
Build and test / Android (aarch64) (push) Failing after 33m33s
They were asked for and they are here, but not on the same footing as the
Kodak profiles, and the files say so in their first line.

What I had claimed, and had to withdraw: that Delta 100 is "quoted around 9"
and HP5 "around 12". Ilford publish no such figures. The word granularity does
not occur anywhere in their technical information -- grain is described as
"fine" and "finest" and nothing more. That claim was in this crate's
documentation as though it came from a datasheet; it is corrected there too.

Two further traps found while looking:

  - Kodak colour negatives publish Print Grain Index, not RMS granularity.
    PGI is a perceptual scale from viewer surveys -- 25 is roughly the
    threshold of visibility, four units a just-noticeable difference -- and
    Kodak state it cannot be compared to RMS. So a Portra number cannot be
    dropped into the granularity field, and none has been.
  - RMS proper is published mostly for black-and-white, reversal and motion
    picture stocks. Every shipped stock therefore still carries the same
    default, which means grain does not yet tell one film from another. That
    is per-stock data, not code, and is now written down where somebody will
    find it.

So the Ilford profiles are built rather than extracted, and each part rests on
something different:

  speed        published and exact -- ISO 400/27 for HP5 is a fact
  contrast     ISO 6:1993's normal development, average gradient 0.62
  spectral     borrowed from Kodak Double-X, a *measured* panchromatic
               negative, shifted by the speed difference. Conventional
               panchromatic sensitisation is much alike across black-and-white
               films, and this is far better founded than reading pixels off a
               printed curve
  silver       neutral, which is not an approximation: developed silver
               absorbs flat, and Double-X's measurement is flat
  granularity  estimated, ordered by each film's known relative grain

They render as a film of that speed and contrast. They are not a measurement
of that emulsion, and the two stocks that share a speed differ only in the
estimated part.

`every_shipped_stock_bakes` is tightened to match, because a constructed
profile fails in a way a measured one does not: the curve parses, bakes, and
sits entirely off one end of its own exposure range, rendering every frame
black or blown while passing a finiteness check. It now asserts mid-grey lands
somewhere photographic and that the tone response runs the way the stock's
kind says it should.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 10:08:51 +02:00
dtourolleandClaude Opus 5 3b5952769b Emit floats an f32 can hold, and drop the format! that formats nothing
CI runs cargo fmt --check and clippy -D warnings, and this branch had
never been through either. Both would have failed it.

The bulk was the generated colour tables: eight significant figures where
an f32 carries about 7.2, so the eighth is noise that rounds away at
compile time and clippy's excessive_precision says so 109 times over.
Fixed in the generator rather than only in the file, so it stays fixed --
and the file is trimmed in place rather than re-derived, because
regenerating it needs a colour-science stack that has nothing to do with
the defect.

The format! in the composer is mine too, from extracting the rendering
tail: the braces in it were escaped because the text used to live inside a
larger template, and once extracted the escapes are noise and the call
formats nothing.

Also here, and clearly not mine: an unused import and a shadowed binding
in dr-gpu, and an unused import in a test. They are pre-existing --
clippy has been failing on master before this branch existed, on lints
like is_multiple_of that arrived with a toolchain rather than with
anyone's code. Fixed because CI cannot go green around them, and called
out because a merge commit is a bad place to quietly edit someone else's
crate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 22:28:14 +02:00
dtourolleandClaude Opus 5 6d18517d28 Ship every stock that exists, black and white included
Three profiles was what the first cut needed to prove the model. This is
the rest of the open data: 23 camera stocks and 9 papers, which is all of
spektrafilm.

Black and white was the gap, and it turned out not to be a gap in the
data -- it was a gap in where I looked. Upstream's `main` has 28 colour
profiles and nothing monochrome; `dev` has three more, and they are
Tri-X, Double-X and the 2302 print film they go onto. So the answer to
"do we have B&W" was yes all along, and it needed the dev branch rather
than a fortnight digitising Ilford's datasheet graphs by eye. Those three
are pinned to `dev` per stock; the colour stocks stay on the released
branch.

A monochrome profile is single-channel -- one emulsion, not three -- and
spreading that one layer across all three is exact rather than an
approximation: three layers with identical sensitivity and identical
curves respond identically, which is what one layer does. The dye is the
trap. The renderer *sums* the three layers' contributions, so replicating
it unchanged renders every frame three times too dense -- neutrally, and
therefore plausibly. A third each reconstructs the single emulsion, and
two tests hold both halves: that the densities stay equal, and that they
sum to one emulsion and not three.

Double-X and 2302 ship five curves apiece, measured at five development
times -- 4 to 12 minutes for Double-X. That is push and pull processing as
measured data. The standard 6.5 minutes is what ships; the rest is in the
upstream file waiting for a control to ask for it.

Two stocks are `support: film` and are nevertheless what a negative is
printed *onto*: the cine projection films 2383 and 2393, which the
Vision3 stocks print to. Filtering the picker on support alone offered a
projection stock as something to load in a camera, so it filters on stage,
with a test saying so.

The picker had to change shape twice over. Chips were right for three
stocks and off the edge of a 280px column at twenty-four, and the column
that replaced them was a thousand pixels standing between the
photographer and every slider below. It is a disclosure now: one row
carrying the answer, opened to change it, closed again on choosing. That
is the opposite of the argument this panel used to take the lids off its
sliders, and deliberately so -- an instrument you compare wants to be
visible, and a list you consult once wants to be out of the way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 20:30:11 +02:00
dtourolleandClaude Opus 5 b6a95e1965 Simulate a film stock from its measurements, not from someone's grade
FR-DEV-3f asks for look emulation and proposes HaldCLUT import to inherit
the free film-simulation ecosystem. This takes the other road for the
stocks where the measurements exist: run the physics.

A stock here is its manufacturer's own datasheet -- spectral sensitivity,
characteristic curves, dye densities. Light exposes three emulsion layers,
the layers develop to densities, the densities are dyes that absorb, and
what is left is what reaches the eye. A colour negative comes out orange
and upside down because that is what a colour negative is; it becomes a
photograph when a paper profile prints it, with the enlarger's filtration
solved rather than dialled.

What that buys over a LUT is that the parameters stay physical. Opening up
a stop moves the picture along the film's real characteristic curve,
shoulder and all, instead of scaling a number baked at one exposure. The
data cost runs the other way too: a stock is 17 kB of published
measurements where one HaldCLUT is 800 kB of one person's grade.

It looks like it needs a spectral integration per pixel. It does not, and
that is the whole design:

  - Exposure is a 3x3 matrix. The reconstructed scene spectrum is linear
    in the sRGB triple, so the integral collapses into nine numbers,
    exactly -- no approximation.
  - The characteristic curve is three 1D functions, sampled exactly.
  - Everything after that -- dye absorption, the print through the
    negative, the paper, the viewing illuminant, the adaptation -- takes
    exactly three numbers in, so it bakes into one 32^3 lookup.

Per pixel: a matrix multiply, three curve taps, one fetch. Splitting the
curve out of the 3D lookup rather than baking one LUT over exposure is
measured, not assumed: the curve carries the sharp shape and the dye
mixing is smooth, so folding them together would need three times the
resolution for the same error. At 32^3 the worst error is 0.003 in linear
sRGB, under one 8-bit code value, and a test says so.

No wgpu dependency, deliberately, and the same isolation argument dr-lens
makes: the model is plain f32 with a documented layout, so every property
worth asserting is asserted on the CPU. Binding it to a texture is dr-gpu's
job and is not done here yet.

The expected values in tests/ came from a Python prototype running against
a different colour-science stack. Agreement to three decimals is evidence
about the model rather than about one implementation of it -- a transposed
matrix or a mispasted observer row would pass every unit test and fail
that one.

Profiles are converted from spektrafilm by Andrea Volpato, CC BY-SA 4.0.
The converter is in the tree and runnable, so what was changed from
upstream is auditable rather than taken on trust; profiles/CHANGELOG.txt
records it, including the one deliberate deviation -- Mallett & Yuksel's
1 kB basis instead of Hanatos's 4 MB table, which costs accuracy at the
gamut edge and saves four megabytes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 14:44:54 +02:00
dtourolleandClaude Sonnet 5 e6b01226eb Let the segmentation export script pick its own input size
Comparing a larger model against a larger input size meant re-exporting at
resolutions other than the shipped 640, and the script only ever wrote that
one number. `IMGSZ` is now a second positional argument, defaulted to 640 so
every existing call is unchanged.

The experiment this was built for found bigger input a net loss on its own
merits — yolo26n-seg and yolo26s-seg at 1280 both lost track of large,
frame-filling subjects (a bus's box shrank and its score nearly halved)
in exchange for catching small or partially-occluded ones tiling already
handles. Nothing shipped from it, but the ability to re-run that comparison
is worth keeping.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-23 13:16:51 +02:00
dtourolleandClaude Opus 5 5a0a9719eb Set the version once, in one place, for every artefact
The workspace read 0.1.0 for two releases, so every desktop binary reported a
version two releases stale. The APK was worse: `AndroidManifest.xml` states no
version at all, so a device showed `versionName=null` and `versionCode=0` while
the library inside the APK knew exactly what it was. A version edited by hand
in several files is a version that is wrong in at least one of them.

`tools/set-version.sh` is now the only thing that sets one. It takes the
version from the latest git tag, or is told, and writes the two files that must
state it before anything is built: the workspace `Cargo.toml`, from which every
crate inherits, and `packaging/PKGBUILD`, which pacman reads before a build
exists. It refreshes `Cargo.lock`, because members appear there by version and
CI builds `--locked`. `--commit` commits the result.

Android is not in that list on purpose. `package.sh` reads the version out of
`Cargo.toml` and hands it to `aapt2 link`, so the APK cannot drift from the
binary it contains — there is no third file to forget. `versionCode` has to be
one increasing integer, which a semantic version is not, so it is packed as
MAJOR*10000 + MINOR*100 + PATCH: ordered the way Android requires, and readable
at a glance.

A version that is not MAJOR.MINOR.PATCH is refused rather than coerced. It is
a contract with whoever reads a bug report, and silently turning "0.4" into
something else is worse than being asked to type it again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 11:09:03 +02:00
dtourolle 0da8271836 Let the model say what a thing is and the watershed say where it ends
Local masking needs to know where an image's regions are. The watershed
spike (S15 arm A) found the boundaries but had no idea what any of them
enclosed; its coarse levels were geometric accidents. This adds the other
half and the thing that joins them.

`core/dr-segment` is where region reasoning now lives — the hierarchy moves
out of `dr-gpu`, which keeps only the pixel passes that are genuinely
shaders. The new crate is device-free and, without its default features,
model-free too: 20 of its tests need neither an adapter nor 11 MB of
weights.

Arm B runs YOLO26n-seg through `ort`. D13 framed inference as a choice
between `ort`'s C++ runtime and the pure-Rust dependency policy; that was a
false choice. `ort`'s `alternative-backend` feature unlinks the C entirely
and `ort-tract` supplies the API from tract, which is pure Rust. Measured
before committing to it: zero unsupported operators, 420 ms for 640x640,
and correct masks on bus.jpg. No NDK problem to solve, so D13's largest
tolerated exception is not needed.

Arm C is `prior.rs`, and it ships because the two arms fail in opposite
directions. Instance membership re-weights the merge saddles, so region
pairs the model believes share an object merge early and pairs straddling
its edge merge late. No boundary moves — only the order in which they
dissolve — which is how the result stays pixel-accurate at every level
while its coarse levels become named things.

Two things the spec assumed that turned out to be false, both recorded in
models/LICENCE.md: there is no usable ADE20K-trained YOLO, so the shipped
vocabulary is COCO's 80 subjects and *stuff* like sky and foliage must come
from arm A; and tract cannot parse a dynamic-shape export, so the graph's
input is fixed and tiling is the only route to more semantic resolution.

Weights are AGPL-3.0, which GPLv3 §13 permits and which makes the combined
work effectively AGPL. Deliberate, not accidental. They live in Git LFS,
and a build script fails with an instruction rather than embedding a
pointer file when the clone lacks them.
2026-08-22 08:39:16 +02:00
dtourolleandClaude Opus 5 7c57f490fe Declare a develop operation in YAML, and generate the rest
An operation was, in the overwhelming majority of cases, four facts: what
its parameters are, what uniforms they compute, what WGSL those uniforms
drive, and where it sits in the chain. Written in Rust those four facts
arrived wrapped in ninety lines of trait implementation — a match on
parameter id to a struct field, another match back, an is_active comparing
each field to its default, a Vec<Uniform> built by hand. All mechanical,
and each one a place to make a silent mistake: a param() arm returning the
wrong field reads perfectly and breaks the sidecar round-trip.

So the four facts are the file now. core/dr-pipeline/ops/<id>.yaml is a
node, build.rs compiles it into the same Operation impl as before, and the
result lands in OUT_DIR — the same reasoning as style.yaml -> theme.slint,
including why it does not land beside the sources it would look exactly
like. Nothing downstream can tell a declared node from a hand-written one:
same &'static OpDescriptor, same fused-shader composition, same sidecar.

Nine nodes moved: exposure, white_balance, contrast, highlights_shadows,
blacks_whites, brilliance, vibrance, saturation, and the shared WGSL
helper registry. Their prose came with them, and so did their tests —
set/expect/expect_active/expect_wgsl in the declaration compile to real
#[test]s, so a node file carries its own proof rather than leaving it
behind in a file that no longer exists.

Two stayed in Rust and say so with `rust:`. The tone curve's neutral is a
relationship between five interpolated points rather than a set of values;
the colour mixer generates thirty-six faceted parameters from twelve
computed hue bands. A schema stretched to cover either would be a worse
language than Rust aimed at one caller. They still declare their position
here, because the chain's *order* is the one thing a reader comes to this
directory to learn, and an order written half in YAML and half in Rust
would be worse than either alone. default_chain() is generated from it.

Uniforms are derived by a small expression language — exp2(exposure),
blacks / 100 * 0.02 — compiled to Rust rather than interpreted, so an
unknown name or a wrong arity is a build error naming the file and the key
and the arithmetic costs nothing at runtime. The build script refuses a
duplicate order, a filename disagreeing with its id, a default outside its
own range, a test value the graph would clamp before the node saw it, a
helper that does not define the function it names, and a declared node
colliding with a file in src/ops.

Verified by adding a scratch node and removing it again: one file, no
other edit, and it joined the chain at its declared order with its test
running. 237 tests pass in dr-pipeline, clippy and fmt clean.

.yaml joins the traceability tool's scanned suffixes, because a node's
Rust now lives in OUT_DIR where a tag could never be linked from the
report. Coverage 47.7% -> 48.3%.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 21:08:16 +02:00
dtourolleandClaude Opus 5 03326242a1 Make the CI checks say what they mean, and format the workspace
Build and test / Desktop (Linux) (push) Successful in 1h23m26s
Build and test / Android (aarch64) (push) Failing after 2s
Build and test / Layer separation (push) Successful in 50s
Traceability / Requirement traces (push) Failing after 1m9s
The Android job's "Verify minimum API level" step has never verified the
minimum API level. It took the first `*.so` anywhere under the target
directory, which is a host proc-macro from debug/deps — an x86-64 object
built by the runner's gcc, whose .comment section cannot mention Android
and so can never contradict the expected value. It now reads the artifact
under the target triple, compares against MIN_API parsed from the
Dockerfile rather than a second copy of the number, and fails on a
mismatch. Both sides are checked non-empty first: two failed parses would
otherwise compare equal and pass, which is the same silent success in a
new costume.

The Android image installs one SDK package per layer and keeps the
output. sdkmanager is a JVM program that aborts when it cannot get memory,
and the single `> /dev/null` step reported that as a bare "exit code 134"
while a retry re-downloaded everything that had already succeeded.

tools/ci-local.sh runs all four jobs — desktop, android, layering,
traceability — against the host toolchain, which is pinned to the same
1.92.0 CI installs. Its matrix check compares regeneration against the
working tree rather than against HEAD: CI starts from a clean checkout, so
git's answer is the right one there and reports every local run stale here.

The rest is rustfmt across the workspace, and the clippy findings that
surfaced once it did: manual_contains in dr-thumbs and collections_ui, a
map iterated as pairs for its keys, an index loop over a slice, and two
runtime assertions on a constant now made at compile time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 12:02:51 +02:00
dtourolle 0f202fd3f9 Add requirements traceability gate and Gitea pipelines
Ports JellyTau's traceability tooling to Rust, carrying across the bug it
was repaired for. That gate divided a traced count by frozen literal
denominators; the requirements file outgrew them and it reported 158%
coverage, so it could never fail its own threshold.

Two rules, both enforced by the extractor's own tests:

  - denominators parsed from docs/requirements.md at run time
  - coverage is |traced ∩ defined| / |defined|, never a raw traced count

The gate additionally fails hard on a misconfigured run — zero
requirements parsed or zero files scanned — rather than reporting a
plausible 0%, and on any orphan tag naming a requirement that does not
exist.

Adapted for DarkRoom: IDs are FR-CAT-1 / NFR-P13 / FR-DEV-3a shapes
rather than JellyTau's fixed three digits, and decisions (D), spikes (S),
milestone items (M) and test ids remain taggable while being excluded
from the denominator — counting them inflated it by 25.

Also adds dr-sync: the RemoteBackend trait and capability model, so the
Nextcloud connector is one implementation rather than the only shape the
engine understands. No mature Nextcloud crate exists (reqwest_dav is too
thin), so the connector will be hand-rolled over reqwest per D7.

Gitea workflows follow the same style: containerised, commented with the
reasoning, desktop and Android on every push, plus a CI check that no
core/ crate depends on the UI toolkit (ARCH §6.5a).

Coverage today: 13.3% (19/143). 50 tests passing.
2026-08-09 08:01:32 +02:00