🐳 Android image / Build and push (push) Successful in 4s
Build and test / android-image (push) Successful in 4s
Build and test / Desktop (Linux) (push) Successful in 1h58m2s
Build and test / Layer separation (push) Successful in 2m3s
Traceability / Requirement traces (push) Successful in 31s
Build and test / Android (aarch64) (push) Failing after 22m34s
The desktop job died mid-link with LLVM reporting "IO failure on output stream", which reads like a compiler crash and is not one: underneath it is `No space left on device`. The runner ran out of disk while linking. Worth knowing what it was spending it on. `target/debug` was 24 GB against `target/release`'s 2.6 GB -- the test build is roughly ninety percent of the footprint -- and of that, 15 GB was debug info in `debug/deps` and 3.6 GB was incremental state. Neither buys anything here. Nothing attaches a debugger to a CI run, and incremental compilation exists to make the second build in a working tree fast, which is not a thing a fresh checkout ever has. With both off the same tree is 3.3 GB, `debug/deps` 2.8 GB, and the test binaries build unchanged. Backtraces keep function names and lose file and line numbers; if a failure ever needs those, DEBUG=1 gives line tables back for a fraction of the 15 GB. A `df -h` either side of the expensive steps, so the next time this happens it says so in one line rather than as an error from LLVM. This is a mitigation and it should not be mistaken for the fix. It bounds what this job asks for; it cannot help if the runner is full of anything else, and 24 GB of build output is not obviously the largest thing on a host that also keeps every cached target directory this workflow has ever saved. If it fails here again, the disk needs looking at on draco-x86. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
354 lines
16 KiB
YAML
354 lines
16 KiB
YAML
name: Build and test
|
|
|
|
# Desktop and Android are built on every push, per the v0.1 decision to carry
|
|
# both platforms from the first commit. An Android break is then caught the day
|
|
# it lands rather than at a porting milestone.
|
|
|
|
on:
|
|
push:
|
|
branches: [main, master, develop]
|
|
pull_request:
|
|
branches: [main, master, develop]
|
|
|
|
jobs:
|
|
# The Android job runs inside an image that this repo builds. Ensure it is in
|
|
# the registry before anything tries to pull it — see android-image.yml for
|
|
# why this is a job rather than a documented manual step. It is a no-op of a
|
|
# few seconds unless docker/android actually changed.
|
|
android-image:
|
|
uses: ./.gitea/workflows/android-image.yml
|
|
|
|
desktop:
|
|
runs-on: linux/amd64
|
|
name: Desktop (Linux)
|
|
# actions/checkout and actions/cache are JavaScript actions: the runner
|
|
# executes them with Node from inside this container. The bare runner image
|
|
# has none, so the job failed at checkout before reaching any build step.
|
|
container:
|
|
image: catthehacker/ubuntu:act-latest
|
|
|
|
# This job filled the runner's disk and died mid-link with "No space left
|
|
# on device" — LLVM reporting an IO failure on its output stream, which
|
|
# reads like a compiler crash and is not one.
|
|
#
|
|
# `target/debug` was 24 GB against `target/release`'s 2.6 GB: 15 GB of it
|
|
# debug info in `debug/deps`, 3.6 GB incremental state. Neither earns its
|
|
# space here. Nothing attaches a debugger to a CI run, and incremental
|
|
# compilation exists to make the *second* build in a working tree fast,
|
|
# which is not a thing a fresh checkout has. Turning both off is the
|
|
# standard CI setting rather than a trick.
|
|
#
|
|
# Measured on this workspace: the same `cargo test --workspace --no-run`
|
|
# tree goes from 24 GB to 3.3 GB, `debug/deps` from 15 GB to 2.8 GB.
|
|
#
|
|
# Backtraces still name functions without debug info; they lose file and
|
|
# line numbers. If a test failure ever needs those, drop DEBUG to 1
|
|
# (line-tables-only) rather than back to 2.
|
|
#
|
|
# This is a mitigation, not a fix. If the runner is full of anything other
|
|
# than this job's own output, it will still be full afterwards.
|
|
env:
|
|
CARGO_INCREMENTAL: 0
|
|
CARGO_PROFILE_DEV_DEBUG: 0
|
|
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@v4
|
|
|
|
# The model, which is in LFS and is not optional.
|
|
#
|
|
# `core/dr-segment/models/*.onnx` is tracked in LFS (.gitattributes), so a
|
|
# plain checkout writes a ~130-byte pointer where 11 MB should be, and
|
|
# `dr-segment`'s build script panics by design rather than embedding a
|
|
# pointer and failing at inference. That failure reads like a broken build
|
|
# instead of a missing fetch, which is how it went unnoticed.
|
|
#
|
|
# Not `lfs: true` on the checkout above, and no `Authorization` header
|
|
# here either. Both install a blanket header for every request to this
|
|
# host, and the object download is the one request that already carries
|
|
# one: `git lfs pull` asks `/info/lfs/objects/batch` first, and Gitea
|
|
# answers with a short-lived `Bearer` JWT scoped to that object. git-lfs
|
|
# then sends the JWT *and* the configured header, and two `Authorization`
|
|
# headers is a 400 from Gitea — reported as
|
|
# LFS: Client error: .../info/lfs/objects/<oid>
|
|
# one step after the batch call that had just succeeded, which reads like
|
|
# a rejected credential rather than a duplicated one. A lone token header
|
|
# is understood fine; it is only the collision that fails.
|
|
#
|
|
# So: strip the headers and hand the token to git-lfs as an ordinary
|
|
# credential instead. It authenticates the batch call and leaves the
|
|
# per-object JWT untouched. Gitea authenticates on the password, so the
|
|
# username is a placeholder. Nothing later in this job talks to the
|
|
# remote, so dropping checkout's header costs us nothing.
|
|
#
|
|
# `continue-on-error` deliberately: if this cannot authenticate, the build
|
|
# below still runs and fails with the build script's own message, which
|
|
# names the real problem. A checkout that dies here says nothing.
|
|
- name: Fetch the segmentation model
|
|
continue-on-error: true
|
|
env:
|
|
LFS_TOKEN: ${{ secrets.GITEA_TOKEN || github.token }}
|
|
run: |
|
|
set -e
|
|
git lfs install --local
|
|
git config --local --get-regexp '^http\..*extraheader$' \
|
|
| cut -d' ' -f1 | sort -u \
|
|
| while read -r key; do git config --local --unset-all "$key"; done || true
|
|
git config --local lfs.url \
|
|
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
|
|
git lfs pull
|
|
ls -l core/dr-segment/models/
|
|
|
|
- name: Cache cargo
|
|
uses: actions/cache@v4
|
|
with:
|
|
path: |
|
|
~/.cargo/registry
|
|
~/.cargo/git
|
|
target
|
|
key: desktop-${{ runner.os }}-${{ hashFiles('**/Cargo.lock') }}
|
|
|
|
# Slint and winit need these at build time; the runner image is minimal.
|
|
- name: Build dependencies
|
|
run: |
|
|
apt-get update -qq
|
|
apt-get install -y -qq pkg-config libfontconfig1-dev libxkbcommon-dev
|
|
|
|
# The act image ships Node but no Rust. Pinned to the workspace
|
|
# rust-version so CI, the Android image, and local builds agree — a
|
|
# floating toolchain turns an unrelated push into a mystery failure.
|
|
- name: Install Rust 1.92.0
|
|
run: |
|
|
set -e
|
|
curl -fsSL https://sh.rustup.rs | sh -s -- \
|
|
-y --no-modify-path --profile minimal \
|
|
--default-toolchain 1.92.0 --component rustfmt,clippy
|
|
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
|
|
|
|
# Free space before and after the expensive steps, so a repeat of the
|
|
# disk exhaustion above is one line to diagnose instead of a puzzling
|
|
# LLVM error.
|
|
- name: Disk before
|
|
run: df -h /workspace 2>/dev/null || df -h .
|
|
|
|
- name: Format
|
|
run: cargo fmt --all -- --check
|
|
|
|
- name: Clippy
|
|
run: cargo clippy --workspace --all-targets -- -D warnings
|
|
|
|
# GPU tests skip themselves where no adapter is present rather than
|
|
# failing — CI runners generally have none, and a test that cannot run is
|
|
# not evidence either way.
|
|
- name: Test
|
|
run: cargo test --workspace
|
|
|
|
- name: Build
|
|
run: cargo build --workspace --release
|
|
|
|
- name: Disk after
|
|
if: always()
|
|
run: df -h /workspace 2>/dev/null || df -h .
|
|
|
|
android:
|
|
runs-on: linux/amd64
|
|
name: Android (aarch64)
|
|
# Waits for the image build. Without this the pull races the push and the
|
|
# job dies with "manifest unknown" before its first step, which is the
|
|
# failure mode this ordering exists to remove.
|
|
needs: android-image
|
|
container:
|
|
image: gitea.tourolle.paris/dtourolle/darkroom-android:latest
|
|
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@v4
|
|
|
|
# The model, which is in LFS and is not optional.
|
|
#
|
|
# `core/dr-segment/models/*.onnx` is tracked in LFS (.gitattributes), so a
|
|
# plain checkout writes a ~130-byte pointer where 11 MB should be, and
|
|
# `dr-segment`'s build script panics by design rather than embedding a
|
|
# pointer and failing at inference. That failure reads like a broken build
|
|
# instead of a missing fetch, which is how it went unnoticed.
|
|
#
|
|
# Not `lfs: true` on the checkout above, and no `Authorization` header
|
|
# here either. Both install a blanket header for every request to this
|
|
# host, and the object download is the one request that already carries
|
|
# one: `git lfs pull` asks `/info/lfs/objects/batch` first, and Gitea
|
|
# answers with a short-lived `Bearer` JWT scoped to that object. git-lfs
|
|
# then sends the JWT *and* the configured header, and two `Authorization`
|
|
# headers is a 400 from Gitea — reported as
|
|
# LFS: Client error: .../info/lfs/objects/<oid>
|
|
# one step after the batch call that had just succeeded, which reads like
|
|
# a rejected credential rather than a duplicated one. A lone token header
|
|
# is understood fine; it is only the collision that fails.
|
|
#
|
|
# So: strip the headers and hand the token to git-lfs as an ordinary
|
|
# credential instead. It authenticates the batch call and leaves the
|
|
# per-object JWT untouched. Gitea authenticates on the password, so the
|
|
# username is a placeholder. Nothing later in this job talks to the
|
|
# remote, so dropping checkout's header costs us nothing.
|
|
#
|
|
# `continue-on-error` deliberately: if this cannot authenticate, the build
|
|
# below still runs and fails with the build script's own message, which
|
|
# names the real problem. A checkout that dies here says nothing.
|
|
- name: Fetch the segmentation model
|
|
continue-on-error: true
|
|
env:
|
|
LFS_TOKEN: ${{ secrets.GITEA_TOKEN || github.token }}
|
|
run: |
|
|
set -e
|
|
git lfs install --local
|
|
git config --local --get-regexp '^http\..*extraheader$' \
|
|
| cut -d' ' -f1 | sort -u \
|
|
| while read -r key; do git config --local --unset-all "$key"; done || true
|
|
git config --local lfs.url \
|
|
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
|
|
git lfs pull
|
|
ls -l core/dr-segment/models/
|
|
|
|
- name: Cache cargo
|
|
uses: actions/cache@v4
|
|
with:
|
|
path: |
|
|
/opt/cargo/registry
|
|
target-android
|
|
key: android-${{ hashFiles('**/Cargo.lock') }}
|
|
|
|
# A fast gate on the crates most likely to break the cross-compile, run
|
|
# before the expensive part. It is `cargo check`, so it type-checks
|
|
# without linking and returns in a fraction of the time the step below
|
|
# takes.
|
|
#
|
|
# Not a statement that only these crates cross-compile — `darkroom-android`
|
|
# and the whole UI stack beneath it build for aarch64 too, which is what
|
|
# the API-level step below does. This one exists to fail fast and name a
|
|
# smaller suspect when it does.
|
|
- name: Cross-compile core
|
|
env:
|
|
CARGO_TARGET_DIR: target-android
|
|
run: cargo check -p dr-types -p dr-gpu -p dr-sync --target aarch64-linux-android
|
|
|
|
# The linker targets MIN_API, not the compile SDK. cargo-ndk otherwise
|
|
# defaults to API 21, far below the Vulkan floor this app needs — and the
|
|
# mismatch is invisible until a device refuses to install.
|
|
#
|
|
# Look under the target triple, and fail on a mismatch. Searching the
|
|
# whole target dir for the first `*.so` found the host proc-macro
|
|
# libraries in target-android/debug/deps instead — x86-64 objects built
|
|
# by the runner's gcc, whose .comment section says nothing about Android
|
|
# and can never contradict the expected API. The step passed regardless
|
|
# of what the linker actually did, which is the one thing it exists to
|
|
# rule out.
|
|
- name: Verify minimum API level
|
|
env:
|
|
CARGO_TARGET_DIR: target-android
|
|
run: |
|
|
set -e
|
|
# `darkroom-android`, not a core crate: this step reads the API level
|
|
# out of a *linked* object, and only that crate produces one. It is
|
|
# the workspace's single `crate-type = ["cdylib"]`; a library crate
|
|
# builds an rlib, which is an archive of object files that no linker
|
|
# has yet touched and that `file` therefore has nothing to say about.
|
|
# Asking for `-p dr-gpu` here could only ever reach the "no aarch64
|
|
# .so was produced" branch below, whatever the linker did.
|
|
#
|
|
# It is also the honest artefact to check: the .so this names is the
|
|
# one that ships in the APK, so the API level verified here is the
|
|
# API level a device will refuse to install against.
|
|
cargo ndk -t arm64-v8a -o target-android/jniLibs \
|
|
build -p darkroom-android --release
|
|
MIN_API=$(sed -n 's/^ARG MIN_API=\([0-9]*\).*/\1/p' docker/android/Dockerfile)
|
|
# Empty on both sides would compare equal and pass, so neither side
|
|
# is allowed to be the result of a failed parse.
|
|
if [ -z "$MIN_API" ]; then
|
|
echo "no ARG MIN_API= in docker/android/Dockerfile"
|
|
exit 1
|
|
fi
|
|
SO=$(find target-android/aarch64-linux-android/release -maxdepth 1 -name '*.so' | head -1)
|
|
if [ -z "$SO" ]; then
|
|
echo "no aarch64 .so was produced"
|
|
exit 1
|
|
fi
|
|
echo "checking $SO"
|
|
file "$SO"
|
|
API=$(file "$SO" | sed -n 's/.*for Android \([0-9]*\).*/\1/p')
|
|
if [ -z "$API" ] || [ "$API" != "$MIN_API" ]; then
|
|
echo "FAIL: linked for Android '${API:-unknown}', expected $MIN_API"
|
|
exit 1
|
|
fi
|
|
|
|
# The APK itself, so a run leaves something installable behind rather
|
|
# than only the knowledge that it would have linked. The assembly is
|
|
# `docker/android/assemble-apk.sh`, shared with `package.sh` so the file
|
|
# a device gets from `package.sh --install` and the file published here
|
|
# are built by the same code — see that script's header.
|
|
#
|
|
# `KEYSTORE` deliberately points at a throwaway directory instead of its
|
|
# default under `target-android`: that directory is what `actions/cache`
|
|
# restores and saves, and a signing key has no business in a build cache
|
|
# or in anything this job uploads. A fresh debug key per run is the right
|
|
# trade for an artefact whose purpose is getting the app onto a test
|
|
# device; nothing upgrades in place over it, which is the one thing a
|
|
# stable key would buy.
|
|
- name: Package the APK
|
|
env:
|
|
CARGO_TARGET_DIR: target-android
|
|
run: |
|
|
set -e
|
|
KEYDIR="$(mktemp -d)"
|
|
trap 'rm -rf "$KEYDIR"' EXIT
|
|
REPO="$PWD" \
|
|
TARGET_DIR="$PWD/target-android" \
|
|
KEYSTORE="$KEYDIR/debug.keystore" \
|
|
bash docker/android/assemble-apk.sh
|
|
|
|
# `if-no-files-found: error` because the failure this guards against is
|
|
# a green run with an empty artefact list, which reads as success until
|
|
# somebody goes looking for the file.
|
|
- name: Upload the APK
|
|
uses: actions/upload-artifact@v4
|
|
with:
|
|
name: darkroom-arm64-v8a-apk
|
|
path: target-android/apk/darkroom.apk
|
|
if-no-files-found: error
|
|
|
|
layering:
|
|
runs-on: linux/amd64
|
|
name: Layer separation
|
|
# Node for the JS actions, as above. cargo comes from rustup below.
|
|
container:
|
|
image: catthehacker/ubuntu:act-latest
|
|
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@v4
|
|
|
|
# `cargo tree` resolves the dependency graph, so it needs the registry
|
|
# index but no system libraries — this job builds nothing.
|
|
- name: Install Rust 1.92.0
|
|
run: |
|
|
set -e
|
|
curl -fsSL https://sh.rustup.rs | sh -s -- \
|
|
-y --no-modify-path --profile minimal --default-toolchain 1.92.0
|
|
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
|
|
|
|
# ARCH §6.5a: no core/ crate may depend on the UI toolkit. One stray
|
|
# `use slint::` costs headless golden-image testing and the
|
|
# one-operation-two-presentations property together, and nothing else
|
|
# would notice.
|
|
- name: Core crates must not depend on the UI
|
|
run: |
|
|
set -e
|
|
FAILED=0
|
|
for crate in dr-types dr-gpu dr-sync; do
|
|
if cargo tree -p "$crate" -e normal 2>/dev/null | grep -qE '\bslint\b|\bi-slint'; then
|
|
echo "FAIL: $crate depends on Slint (ARCH §6.5a)"
|
|
FAILED=1
|
|
else
|
|
echo "ok: $crate"
|
|
fi
|
|
done
|
|
exit $FAILED
|