Files
DarkRoom/.gitea/workflows/build-and-test.yml
T
dtourolleandClaude Opus 5 56bdd457dd
🐳 Android image / Build and push (push) Successful in 4s
Build and test / android-image (push) Successful in 4s
Build and test / Desktop (Linux) (push) Successful in 1h58m2s
Build and test / Layer separation (push) Successful in 2m3s
Traceability / Requirement traces (push) Successful in 31s
Build and test / Android (aarch64) (push) Failing after 22m34s
Stop the test build filling the runner's disk
The desktop job died mid-link with LLVM reporting "IO failure on output
stream", which reads like a compiler crash and is not one: underneath it
is `No space left on device`. The runner ran out of disk while linking.

Worth knowing what it was spending it on. `target/debug` was 24 GB
against `target/release`'s 2.6 GB -- the test build is roughly ninety
percent of the footprint -- and of that, 15 GB was debug info in
`debug/deps` and 3.6 GB was incremental state. Neither buys anything
here. Nothing attaches a debugger to a CI run, and incremental
compilation exists to make the second build in a working tree fast,
which is not a thing a fresh checkout ever has.

With both off the same tree is 3.3 GB, `debug/deps` 2.8 GB, and the test
binaries build unchanged. Backtraces keep function names and lose file
and line numbers; if a failure ever needs those, DEBUG=1 gives line
tables back for a fraction of the 15 GB.

A `df -h` either side of the expensive steps, so the next time this
happens it says so in one line rather than as an error from LLVM.

This is a mitigation and it should not be mistaken for the fix. It bounds
what this job asks for; it cannot help if the runner is full of anything
else, and 24 GB of build output is not obviously the largest thing on a
host that also keeps every cached target directory this workflow has ever
saved. If it fails here again, the disk needs looking at on draco-x86.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 01:13:32 +02:00

354 lines
16 KiB
YAML

name: Build and test
# Desktop and Android are built on every push, per the v0.1 decision to carry
# both platforms from the first commit. An Android break is then caught the day
# it lands rather than at a porting milestone.
on:
push:
branches: [main, master, develop]
pull_request:
branches: [main, master, develop]
jobs:
# The Android job runs inside an image that this repo builds. Ensure it is in
# the registry before anything tries to pull it — see android-image.yml for
# why this is a job rather than a documented manual step. It is a no-op of a
# few seconds unless docker/android actually changed.
android-image:
uses: ./.gitea/workflows/android-image.yml
desktop:
runs-on: linux/amd64
name: Desktop (Linux)
# actions/checkout and actions/cache are JavaScript actions: the runner
# executes them with Node from inside this container. The bare runner image
# has none, so the job failed at checkout before reaching any build step.
container:
image: catthehacker/ubuntu:act-latest
# This job filled the runner's disk and died mid-link with "No space left
# on device" — LLVM reporting an IO failure on its output stream, which
# reads like a compiler crash and is not one.
#
# `target/debug` was 24 GB against `target/release`'s 2.6 GB: 15 GB of it
# debug info in `debug/deps`, 3.6 GB incremental state. Neither earns its
# space here. Nothing attaches a debugger to a CI run, and incremental
# compilation exists to make the *second* build in a working tree fast,
# which is not a thing a fresh checkout has. Turning both off is the
# standard CI setting rather than a trick.
#
# Measured on this workspace: the same `cargo test --workspace --no-run`
# tree goes from 24 GB to 3.3 GB, `debug/deps` from 15 GB to 2.8 GB.
#
# Backtraces still name functions without debug info; they lose file and
# line numbers. If a test failure ever needs those, drop DEBUG to 1
# (line-tables-only) rather than back to 2.
#
# This is a mitigation, not a fix. If the runner is full of anything other
# than this job's own output, it will still be full afterwards.
env:
CARGO_INCREMENTAL: 0
CARGO_PROFILE_DEV_DEBUG: 0
steps:
- name: Checkout
uses: actions/checkout@v4
# The model, which is in LFS and is not optional.
#
# `core/dr-segment/models/*.onnx` is tracked in LFS (.gitattributes), so a
# plain checkout writes a ~130-byte pointer where 11 MB should be, and
# `dr-segment`'s build script panics by design rather than embedding a
# pointer and failing at inference. That failure reads like a broken build
# instead of a missing fetch, which is how it went unnoticed.
#
# Not `lfs: true` on the checkout above, and no `Authorization` header
# here either. Both install a blanket header for every request to this
# host, and the object download is the one request that already carries
# one: `git lfs pull` asks `/info/lfs/objects/batch` first, and Gitea
# answers with a short-lived `Bearer` JWT scoped to that object. git-lfs
# then sends the JWT *and* the configured header, and two `Authorization`
# headers is a 400 from Gitea — reported as
# LFS: Client error: .../info/lfs/objects/<oid>
# one step after the batch call that had just succeeded, which reads like
# a rejected credential rather than a duplicated one. A lone token header
# is understood fine; it is only the collision that fails.
#
# So: strip the headers and hand the token to git-lfs as an ordinary
# credential instead. It authenticates the batch call and leaves the
# per-object JWT untouched. Gitea authenticates on the password, so the
# username is a placeholder. Nothing later in this job talks to the
# remote, so dropping checkout's header costs us nothing.
#
# `continue-on-error` deliberately: if this cannot authenticate, the build
# below still runs and fails with the build script's own message, which
# names the real problem. A checkout that dies here says nothing.
- name: Fetch the segmentation model
continue-on-error: true
env:
LFS_TOKEN: ${{ secrets.GITEA_TOKEN || github.token }}
run: |
set -e
git lfs install --local
git config --local --get-regexp '^http\..*extraheader$' \
| cut -d' ' -f1 | sort -u \
| while read -r key; do git config --local --unset-all "$key"; done || true
git config --local lfs.url \
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
git lfs pull
ls -l core/dr-segment/models/
- name: Cache cargo
uses: actions/cache@v4
with:
path: |
~/.cargo/registry
~/.cargo/git
target
key: desktop-${{ runner.os }}-${{ hashFiles('**/Cargo.lock') }}
# Slint and winit need these at build time; the runner image is minimal.
- name: Build dependencies
run: |
apt-get update -qq
apt-get install -y -qq pkg-config libfontconfig1-dev libxkbcommon-dev
# The act image ships Node but no Rust. Pinned to the workspace
# rust-version so CI, the Android image, and local builds agree — a
# floating toolchain turns an unrelated push into a mystery failure.
- name: Install Rust 1.92.0
run: |
set -e
curl -fsSL https://sh.rustup.rs | sh -s -- \
-y --no-modify-path --profile minimal \
--default-toolchain 1.92.0 --component rustfmt,clippy
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
# Free space before and after the expensive steps, so a repeat of the
# disk exhaustion above is one line to diagnose instead of a puzzling
# LLVM error.
- name: Disk before
run: df -h /workspace 2>/dev/null || df -h .
- name: Format
run: cargo fmt --all -- --check
- name: Clippy
run: cargo clippy --workspace --all-targets -- -D warnings
# GPU tests skip themselves where no adapter is present rather than
# failing — CI runners generally have none, and a test that cannot run is
# not evidence either way.
- name: Test
run: cargo test --workspace
- name: Build
run: cargo build --workspace --release
- name: Disk after
if: always()
run: df -h /workspace 2>/dev/null || df -h .
android:
runs-on: linux/amd64
name: Android (aarch64)
# Waits for the image build. Without this the pull races the push and the
# job dies with "manifest unknown" before its first step, which is the
# failure mode this ordering exists to remove.
needs: android-image
container:
image: gitea.tourolle.paris/dtourolle/darkroom-android:latest
steps:
- name: Checkout
uses: actions/checkout@v4
# The model, which is in LFS and is not optional.
#
# `core/dr-segment/models/*.onnx` is tracked in LFS (.gitattributes), so a
# plain checkout writes a ~130-byte pointer where 11 MB should be, and
# `dr-segment`'s build script panics by design rather than embedding a
# pointer and failing at inference. That failure reads like a broken build
# instead of a missing fetch, which is how it went unnoticed.
#
# Not `lfs: true` on the checkout above, and no `Authorization` header
# here either. Both install a blanket header for every request to this
# host, and the object download is the one request that already carries
# one: `git lfs pull` asks `/info/lfs/objects/batch` first, and Gitea
# answers with a short-lived `Bearer` JWT scoped to that object. git-lfs
# then sends the JWT *and* the configured header, and two `Authorization`
# headers is a 400 from Gitea — reported as
# LFS: Client error: .../info/lfs/objects/<oid>
# one step after the batch call that had just succeeded, which reads like
# a rejected credential rather than a duplicated one. A lone token header
# is understood fine; it is only the collision that fails.
#
# So: strip the headers and hand the token to git-lfs as an ordinary
# credential instead. It authenticates the batch call and leaves the
# per-object JWT untouched. Gitea authenticates on the password, so the
# username is a placeholder. Nothing later in this job talks to the
# remote, so dropping checkout's header costs us nothing.
#
# `continue-on-error` deliberately: if this cannot authenticate, the build
# below still runs and fails with the build script's own message, which
# names the real problem. A checkout that dies here says nothing.
- name: Fetch the segmentation model
continue-on-error: true
env:
LFS_TOKEN: ${{ secrets.GITEA_TOKEN || github.token }}
run: |
set -e
git lfs install --local
git config --local --get-regexp '^http\..*extraheader$' \
| cut -d' ' -f1 | sort -u \
| while read -r key; do git config --local --unset-all "$key"; done || true
git config --local lfs.url \
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
git lfs pull
ls -l core/dr-segment/models/
- name: Cache cargo
uses: actions/cache@v4
with:
path: |
/opt/cargo/registry
target-android
key: android-${{ hashFiles('**/Cargo.lock') }}
# A fast gate on the crates most likely to break the cross-compile, run
# before the expensive part. It is `cargo check`, so it type-checks
# without linking and returns in a fraction of the time the step below
# takes.
#
# Not a statement that only these crates cross-compile — `darkroom-android`
# and the whole UI stack beneath it build for aarch64 too, which is what
# the API-level step below does. This one exists to fail fast and name a
# smaller suspect when it does.
- name: Cross-compile core
env:
CARGO_TARGET_DIR: target-android
run: cargo check -p dr-types -p dr-gpu -p dr-sync --target aarch64-linux-android
# The linker targets MIN_API, not the compile SDK. cargo-ndk otherwise
# defaults to API 21, far below the Vulkan floor this app needs — and the
# mismatch is invisible until a device refuses to install.
#
# Look under the target triple, and fail on a mismatch. Searching the
# whole target dir for the first `*.so` found the host proc-macro
# libraries in target-android/debug/deps instead — x86-64 objects built
# by the runner's gcc, whose .comment section says nothing about Android
# and can never contradict the expected API. The step passed regardless
# of what the linker actually did, which is the one thing it exists to
# rule out.
- name: Verify minimum API level
env:
CARGO_TARGET_DIR: target-android
run: |
set -e
# `darkroom-android`, not a core crate: this step reads the API level
# out of a *linked* object, and only that crate produces one. It is
# the workspace's single `crate-type = ["cdylib"]`; a library crate
# builds an rlib, which is an archive of object files that no linker
# has yet touched and that `file` therefore has nothing to say about.
# Asking for `-p dr-gpu` here could only ever reach the "no aarch64
# .so was produced" branch below, whatever the linker did.
#
# It is also the honest artefact to check: the .so this names is the
# one that ships in the APK, so the API level verified here is the
# API level a device will refuse to install against.
cargo ndk -t arm64-v8a -o target-android/jniLibs \
build -p darkroom-android --release
MIN_API=$(sed -n 's/^ARG MIN_API=\([0-9]*\).*/\1/p' docker/android/Dockerfile)
# Empty on both sides would compare equal and pass, so neither side
# is allowed to be the result of a failed parse.
if [ -z "$MIN_API" ]; then
echo "no ARG MIN_API= in docker/android/Dockerfile"
exit 1
fi
SO=$(find target-android/aarch64-linux-android/release -maxdepth 1 -name '*.so' | head -1)
if [ -z "$SO" ]; then
echo "no aarch64 .so was produced"
exit 1
fi
echo "checking $SO"
file "$SO"
API=$(file "$SO" | sed -n 's/.*for Android \([0-9]*\).*/\1/p')
if [ -z "$API" ] || [ "$API" != "$MIN_API" ]; then
echo "FAIL: linked for Android '${API:-unknown}', expected $MIN_API"
exit 1
fi
# The APK itself, so a run leaves something installable behind rather
# than only the knowledge that it would have linked. The assembly is
# `docker/android/assemble-apk.sh`, shared with `package.sh` so the file
# a device gets from `package.sh --install` and the file published here
# are built by the same code — see that script's header.
#
# `KEYSTORE` deliberately points at a throwaway directory instead of its
# default under `target-android`: that directory is what `actions/cache`
# restores and saves, and a signing key has no business in a build cache
# or in anything this job uploads. A fresh debug key per run is the right
# trade for an artefact whose purpose is getting the app onto a test
# device; nothing upgrades in place over it, which is the one thing a
# stable key would buy.
- name: Package the APK
env:
CARGO_TARGET_DIR: target-android
run: |
set -e
KEYDIR="$(mktemp -d)"
trap 'rm -rf "$KEYDIR"' EXIT
REPO="$PWD" \
TARGET_DIR="$PWD/target-android" \
KEYSTORE="$KEYDIR/debug.keystore" \
bash docker/android/assemble-apk.sh
# `if-no-files-found: error` because the failure this guards against is
# a green run with an empty artefact list, which reads as success until
# somebody goes looking for the file.
- name: Upload the APK
uses: actions/upload-artifact@v4
with:
name: darkroom-arm64-v8a-apk
path: target-android/apk/darkroom.apk
if-no-files-found: error
layering:
runs-on: linux/amd64
name: Layer separation
# Node for the JS actions, as above. cargo comes from rustup below.
container:
image: catthehacker/ubuntu:act-latest
steps:
- name: Checkout
uses: actions/checkout@v4
# `cargo tree` resolves the dependency graph, so it needs the registry
# index but no system libraries — this job builds nothing.
- name: Install Rust 1.92.0
run: |
set -e
curl -fsSL https://sh.rustup.rs | sh -s -- \
-y --no-modify-path --profile minimal --default-toolchain 1.92.0
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
# ARCH §6.5a: no core/ crate may depend on the UI toolkit. One stray
# `use slint::` costs headless golden-image testing and the
# one-operation-two-presentations property together, and nothing else
# would notice.
- name: Core crates must not depend on the UI
run: |
set -e
FAILED=0
for crate in dr-types dr-gpu dr-sync; do
if cargo tree -p "$crate" -e normal 2>/dev/null | grep -qE '\bslint\b|\bi-slint'; then
echo "FAIL: $crate depends on Slint (ARCH §6.5a)"
FAILED=1
else
echo "ok: $crate"
fi
done
exit $FAILED