f82c69bc6b3ad2bf7adac89a50e182384940e2e0
12
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
40e6334bb1 |
Sign the APK with a real key when one is configured
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Desktop (Linux) (push) Successful in 21m23s
Build and test / Layer separation (push) Successful in 29s
Traceability / Requirement traces (push) Failing after 28s
Build and test / Android (aarch64) (push) Failing after 33m28s
The APK has been debug-signed with a key generated on the spot, which is right for putting a build on a test device and useless for anything else: a different signature every run, so nothing can ever update in place. Four secrets now select a real signature -- ANDROID_KEYSTORE_BASE64 and its password, alias and key password. The names are JellyTau's, because that repo already signs its Android build this way against this same runner and one convention across both is one thing to remember. Absence of the secrets is not an error. A fork or a branch build has no access to them and should still produce an installable APK, so the debug path stays exactly as it was. The reverse is an error: if a keystore is supplied and cannot be read, the build fails rather than quietly falling back to a debug key, because a release that is silently debug-signed is worse than no release. Passwords reach apksigner and keytool as `env:`, never `pass:`. `pass:` puts the password in the process table for anything on the box to read. The keystore is written to a 0700 mktemp directory and never into the workspace, which is both what actions/cache saves and what the upload step globs. Also: upload-artifact drops from v4 to v3. v4 was a guess about what this Gitea supports. v3 is what JellyTau uploads its APK with on this runner today, which makes it the version known to work rather than the one that ought to. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
56bdd457dd |
Stop the test build filling the runner's disk
🐳 Android image / Build and push (push) Successful in 4s
Build and test / android-image (push) Successful in 4s
Build and test / Desktop (Linux) (push) Successful in 1h58m2s
Build and test / Layer separation (push) Successful in 2m3s
Traceability / Requirement traces (push) Successful in 31s
Build and test / Android (aarch64) (push) Failing after 22m34s
The desktop job died mid-link with LLVM reporting "IO failure on output stream", which reads like a compiler crash and is not one: underneath it is `No space left on device`. The runner ran out of disk while linking. Worth knowing what it was spending it on. `target/debug` was 24 GB against `target/release`'s 2.6 GB -- the test build is roughly ninety percent of the footprint -- and of that, 15 GB was debug info in `debug/deps` and 3.6 GB was incremental state. Neither buys anything here. Nothing attaches a debugger to a CI run, and incremental compilation exists to make the second build in a working tree fast, which is not a thing a fresh checkout ever has. With both off the same tree is 3.3 GB, `debug/deps` 2.8 GB, and the test binaries build unchanged. Backtraces keep function names and lose file and line numbers; if a failure ever needs those, DEBUG=1 gives line tables back for a fraction of the 15 GB. A `df -h` either side of the expensive steps, so the next time this happens it says so in one line rather than as an error from LLVM. This is a mitigation and it should not be mistaken for the fix. It bounds what this job asks for; it cannot help if the runner is full of anything else, and 24 GB of build output is not obviously the largest thing on a host that also keeps every cached target directory this workflow has ever saved. If it fails here again, the disk needs looking at on draco-x86. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
834b219c3f |
Publish the Android APK as a build artefact
Build and test / Desktop (Linux) (push) Failing after 1h12m6s
Build and test / Layer separation (push) Successful in 41s
Traceability / Requirement traces (push) Successful in 23s
🐳 Android image / Build and push (push) Successful in 10m59s
Build and test / android-image (push) Successful in 11m2s
Build and test / Android (aarch64) (push) Failing after 22m52s
The android job proved the app links for aarch64 and then threw the result away. There was no APK anywhere in CI and no upload step in the repo at all, so a green run left nothing anybody could install -- the artefact list was empty by construction, not by failure. It now assembles the APK with the script package.sh uses and uploads it. The .so comes from the build the API-level check already ran; -o only adds a copy of it where the packaging step looks, so this costs one copy rather than a second twenty-minute cross-compile. The signing key is the part worth being careful about. KEYSTORE points at a mktemp directory rather than its default under target-android, because that directory is precisely what actions/cache saves and restores -- the default would have written a private key into the build cache and kept it there. Nothing but the .apk is uploaded. A fresh debug key each run is the right trade for an artefact meant to reach a test device: the only thing a stable key buys is installing over a previous build without uninstalling first, and a key that survives in cache storage to buy it is a bad exchange. if-no-files-found: error because the failure being guarded against is a green run with an empty artefact list, which reads as success right up until somebody goes looking for the file. Debug-signed, arm64-v8a only -- the ABI the job already builds. Neither is a release story; this is a build you can install. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
606f85df34 |
Send the LFS object endpoint one Authorization header, not two
🐳 Android image / Build and push (push) Successful in 2s
Build and test / android-image (push) Successful in 3s
Build and test / Desktop (Linux) (push) Failing after 1h12m2s
Build and test / Layer separation (push) Successful in 36s
Traceability / Requirement traces (push) Failing after 31s
Build and test / Android (aarch64) (push) Failing after 36m17s
The model fetch has been failing on every run with LFS: Client error: .../info/lfs/objects/0672d7a7... which reads like a rejected credential and is not one. The object is on the server and downloads fine; what fails is the shape of the request. `git lfs pull` makes two calls. The first, to `/info/lfs/objects/batch`, succeeds -- and Gitea answers it with a short-lived `Bearer` JWT scoped to that one object, for git-lfs to use on the second. git-lfs sends that JWT *and* the `Authorization` header this step had installed in git config, and two `Authorization` headers is a 400 from Gitea. Hence a client error on the object one step after the batch call it just made successfully, which is what made this look like an auth problem rather than a duplication. Confirmed directly against the server: the JWT alone on that URL is a 200, the JWT plus any second `Authorization` is a 400, and a lone token header that is merely wrong is a 401 -- so the scheme was never the issue. `lfs: true` on the checkout fails the same way and for the same reason, because actions/checkout persists a header of its own; the comment here blaming a credential the endpoint would not accept was wrong on both counts. So the headers are stripped -- checkout's included, since nothing later in either job talks to the remote -- and the token is handed to git-lfs as an ordinary credential instead. It authenticates the batch call and leaves the per-object JWT alone. This is what fails the Android job today: the build script sees a 133-byte pointer and panics by design, which is the message it is supposed to give and the one nobody could act on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
33847a0bbc |
Open one GPU device for the tests, and stop the checkout dying over LFS
Build and test / Desktop (Linux) (push) Failing after 9m18s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 25s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 22m46s
Two CI failures, unrelated except that both were mine. **The test binary faulted under parallel threads.** Every GPU test opened its own `GpuContext`, and `cargo test` runs on as many threads as there are cores — so a full run asked the driver to bring up a dozen Vulkan devices at once and died with SIGSEGV. Serially it passed, which made it look like flakiness rather than a fault in the harness. One device now, behind a `OnceLock`: a `GpuContext` is an `Arc<Device>` and an `Arc<Queue>`, so sharing it is a refcount, and the losers of the race block until the winner is done. 416 tests now pass in parallel, in half the time twenty-six devices took. **`lfs: true` on the checkout made the checkout fail.** The intent was right — the model is in LFS, a plain checkout writes a 133-byte pointer, and the build script panics on it — but on this server `git lfs fetch` is rejected at `/info/lfs/objects/<oid>` with a client error: the credential `actions/checkout` installs for git is not one the LFS endpoint accepts. So a fetch problem presented as a checkout problem and took the whole job with it. The object is on the server; a clean clone over SSH with `git lfs install --local` pulls all 11 MB of it. It is now its own step with an explicit token, and `continue-on-error` so a credential problem cannot masquerade as a broken checkout — if it fails, the build still runs and fails with the build script's own message, which names the real problem. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ec8582032e |
Fetch the model CI has been building without
Both failing jobs failed for the same reason, and it was not the reason either of them appeared to fail. Desktop reported a Clippy failure and Android reported the API-level check; both were `dr-segment`'s build script panicking because `models/yolo26n-seg.onnx` was a 133-byte LFS pointer rather than the 11 MB model. No checkout step asked for LFS objects. `.gitattributes` predicted this exactly — "a clone without git-lfs gets a ~130-byte pointer file where the model should be" — and the build script fails loudly by design rather than embedding a pointer and failing at inference. That design worked; nothing was reading the message. It stayed hidden because the Android job built `-p dr-gpu`, which never reaches dr-segment. Building the app crate does, which is how one fix surfaced another. Only the two jobs that compile get `lfs: true`. `layering` runs `cargo tree` and builds nothing, so it has no reason to fetch 11 MB. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1297e3259b |
Stop claiming the app crates cannot cross-compile
The note above the core-crate check said the UI and app crates would join "once the Android shell exists". They already had. Building `darkroom-android` for aarch64 takes 4m23s and produces `libdarkroom.so`, linked for Android 28 — which is what the step below now checks, and what a device installs. Rewritten to say what the core check is actually for: a fast, link-free gate that fails early and names a smaller suspect, ahead of the full build. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
83528c068a |
Check the API level of an object a linker actually produced
The step built `-p dr-gpu` and then looked for a `*.so` to read the API level out of with `file`. dr-gpu declares no `crate-type`, so it produces an rlib — an archive of object files no linker has touched, and about which `file` has nothing to say. The step could therefore only ever reach its own "no aarch64 .so was produced" branch, whatever the linker did, which is the one thing it exists to rule out. Builds `darkroom-android` instead: the workspace's only `crate-type = ["cdylib"]`, and the artefact that ships in the APK. The API level verified is now the one a device will refuse to install against. The neighbouring comment claiming only core crates cross-compile is left alone until the build proves otherwise; it is either stale or this commit is wrong, and the same run answers both. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b044a8c067 |
Build the Android CI image in CI, not on a laptop
The Android job ran in gitea.tourolle.paris/dtourolle/darkroom-android:latest, a tag that had never been pushed. The image existed only as a local darkroom-android:latest on one machine, so every Android job died at docker pull with "manifest unknown" before reaching a step. The registry API confirms it: that manifest is a 404 while the other builder images answer 200. android-image.yml now builds and pushes it, following KPN's docker.yaml — host runner rather than a container, so it has the Docker daemon and the host's cached registry credentials, and a plain-git checkout because that host has no Node for actions/checkout. Where it diverges from KPN: that workflow gates on dorny/paths-filter running inside the builder image, which here would need the very image that is missing. The tag is the git tree hash of docker/android instead, which changes when and only when a file there changes. An unrelated push reuses the image, a Dockerfile edit cannot keep serving a stale latest, and a missing tag rebuilds itself without a manual step. The presence probe is curl against the registry API, not `docker manifest inspect`. The latter exits 1 on this registry even for tags that are plainly there — jellytau-builder:latest answers HTTP 200 while docker reports "manifest unknown" for it — and trusting it would have rebuilt 7 GB on every push. The HEAD request also yields Docker-Content-Digest, so latest is repointed only when the digests actually disagree, without pulling any layers. A probe that cannot authenticate falls through to building. Rebuilding when it was unnecessary costs minutes; skipping a build that was needed is the failure this commit exists to remove. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
03326242a1 |
Make the CI checks say what they mean, and format the workspace
The Android job's "Verify minimum API level" step has never verified the minimum API level. It took the first `*.so` anywhere under the target directory, which is a host proc-macro from debug/deps — an x86-64 object built by the runner's gcc, whose .comment section cannot mention Android and so can never contradict the expected value. It now reads the artifact under the target triple, compares against MIN_API parsed from the Dockerfile rather than a second copy of the number, and fails on a mismatch. Both sides are checked non-empty first: two failed parses would otherwise compare equal and pass, which is the same silent success in a new costume. The Android image installs one SDK package per layer and keeps the output. sdkmanager is a JVM program that aborts when it cannot get memory, and the single `> /dev/null` step reported that as a bare "exit code 134" while a retry re-downloaded everything that had already succeeded. tools/ci-local.sh runs all four jobs — desktop, android, layering, traceability — against the host toolchain, which is pinned to the same 1.92.0 CI installs. Its matrix check compares regeneration against the working tree rather than against HEAD: CI starts from a clean checkout, so git's answer is the right one there and reports every local run stale here. The rest is rustfmt across the workspace, and the clippy findings that surfaced once it did: manual_contains in dr-thumbs and collections_ui, a map iterated as pairs for its keys, an index loop over a slice, and two runtime assertions on a constant now made at compile time. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4d78041d1d | Many imorovments | ||
|
|
0f202fd3f9 |
Add requirements traceability gate and Gitea pipelines
Ports JellyTau's traceability tooling to Rust, carrying across the bug it was repaired for. That gate divided a traced count by frozen literal denominators; the requirements file outgrew them and it reported 158% coverage, so it could never fail its own threshold. Two rules, both enforced by the extractor's own tests: - denominators parsed from docs/requirements.md at run time - coverage is |traced ∩ defined| / |defined|, never a raw traced count The gate additionally fails hard on a misconfigured run — zero requirements parsed or zero files scanned — rather than reporting a plausible 0%, and on any orphan tag naming a requirement that does not exist. Adapted for DarkRoom: IDs are FR-CAT-1 / NFR-P13 / FR-DEV-3a shapes rather than JellyTau's fixed three digits, and decisions (D), spikes (S), milestone items (M) and test ids remain taggable while being excluded from the denominator — counting them inflated it by 25. Also adds dr-sync: the RemoteBackend trait and capability model, so the Nextcloud connector is one implementation rather than the only shape the engine understands. No mature Nextcloud crate exists (reqwest_dav is too thin), so the connector will be hand-rolled over reqwest per D7. Gitea workflows follow the same style: containerised, commented with the reasoning, desktop and Android on every push, plus a CI check that no core/ crate depends on the UI toolkit (ARCH §6.5a). Coverage today: 13.3% (19/143). 50 tests passing. |