Build and test / Desktop (Linux) (push) Failing after 9m18s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 25s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 22m46s
Two CI failures, unrelated except that both were mine. **The test binary faulted under parallel threads.** Every GPU test opened its own `GpuContext`, and `cargo test` runs on as many threads as there are cores — so a full run asked the driver to bring up a dozen Vulkan devices at once and died with SIGSEGV. Serially it passed, which made it look like flakiness rather than a fault in the harness. One device now, behind a `OnceLock`: a `GpuContext` is an `Arc<Device>` and an `Arc<Queue>`, so sharing it is a refcount, and the losers of the race block until the winner is done. 416 tests now pass in parallel, in half the time twenty-six devices took. **`lfs: true` on the checkout made the checkout fail.** The intent was right — the model is in LFS, a plain checkout writes a 133-byte pointer, and the build script panics on it — but on this server `git lfs fetch` is rejected at `/info/lfs/objects/<oid>` with a client error: the credential `actions/checkout` installs for git is not one the LFS endpoint accepts. So a fetch problem presented as a checkout problem and took the whole job with it. The object is on the server; a clean clone over SSH with `git lfs install --local` pulls all 11 MB of it. It is now its own step with an explicit token, and `continue-on-error` so a credential problem cannot masquerade as a broken checkout — if it fails, the build still runs and fails with the build script's own message, which names the real problem. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
254 lines
11 KiB
YAML
254 lines
11 KiB
YAML
name: Build and test
|
|
|
|
# Desktop and Android are built on every push, per the v0.1 decision to carry
|
|
# both platforms from the first commit. An Android break is then caught the day
|
|
# it lands rather than at a porting milestone.
|
|
|
|
on:
|
|
push:
|
|
branches: [main, master, develop]
|
|
pull_request:
|
|
branches: [main, master, develop]
|
|
|
|
jobs:
|
|
# The Android job runs inside an image that this repo builds. Ensure it is in
|
|
# the registry before anything tries to pull it — see android-image.yml for
|
|
# why this is a job rather than a documented manual step. It is a no-op of a
|
|
# few seconds unless docker/android actually changed.
|
|
android-image:
|
|
uses: ./.gitea/workflows/android-image.yml
|
|
|
|
desktop:
|
|
runs-on: linux/amd64
|
|
name: Desktop (Linux)
|
|
# actions/checkout and actions/cache are JavaScript actions: the runner
|
|
# executes them with Node from inside this container. The bare runner image
|
|
# has none, so the job failed at checkout before reaching any build step.
|
|
container:
|
|
image: catthehacker/ubuntu:act-latest
|
|
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@v4
|
|
|
|
# The model, which is in LFS and is not optional.
|
|
#
|
|
# `core/dr-segment/models/*.onnx` is tracked in LFS (.gitattributes), so a
|
|
# plain checkout writes a ~130-byte pointer where 11 MB should be, and
|
|
# `dr-segment`'s build script panics by design rather than embedding a
|
|
# pointer and failing at inference. That failure reads like a broken build
|
|
# instead of a missing fetch, which is how it went unnoticed.
|
|
#
|
|
# Not `lfs: true` on the checkout above: that makes the *checkout* fail on
|
|
# this server — `git lfs fetch` is rejected at
|
|
# `/info/lfs/objects/<oid>` with a client error, because the credential
|
|
# checkout installs for git is not one the LFS endpoint accepts. Fetching
|
|
# it as its own step with an explicit token keeps a credential problem
|
|
# from looking like a checkout problem, and lets the rest of the job say
|
|
# what it thinks.
|
|
#
|
|
# `continue-on-error` deliberately: if this cannot authenticate, the build
|
|
# below still runs and fails with the build script's own message, which
|
|
# names the real problem. A checkout that dies here says nothing.
|
|
- name: Fetch the segmentation model
|
|
continue-on-error: true
|
|
run: |
|
|
set -x
|
|
git lfs install --local
|
|
git config --local lfs.https://gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs.access basic
|
|
git config --local http.extraheader "Authorization: token ${{ secrets.GITEA_TOKEN || github.token }}"
|
|
git lfs pull
|
|
ls -l core/dr-segment/models/
|
|
|
|
- name: Cache cargo
|
|
uses: actions/cache@v4
|
|
with:
|
|
path: |
|
|
~/.cargo/registry
|
|
~/.cargo/git
|
|
target
|
|
key: desktop-${{ runner.os }}-${{ hashFiles('**/Cargo.lock') }}
|
|
|
|
# Slint and winit need these at build time; the runner image is minimal.
|
|
- name: Build dependencies
|
|
run: |
|
|
apt-get update -qq
|
|
apt-get install -y -qq pkg-config libfontconfig1-dev libxkbcommon-dev
|
|
|
|
# The act image ships Node but no Rust. Pinned to the workspace
|
|
# rust-version so CI, the Android image, and local builds agree — a
|
|
# floating toolchain turns an unrelated push into a mystery failure.
|
|
- name: Install Rust 1.92.0
|
|
run: |
|
|
set -e
|
|
curl -fsSL https://sh.rustup.rs | sh -s -- \
|
|
-y --no-modify-path --profile minimal \
|
|
--default-toolchain 1.92.0 --component rustfmt,clippy
|
|
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
|
|
|
|
- name: Format
|
|
run: cargo fmt --all -- --check
|
|
|
|
- name: Clippy
|
|
run: cargo clippy --workspace --all-targets -- -D warnings
|
|
|
|
# GPU tests skip themselves where no adapter is present rather than
|
|
# failing — CI runners generally have none, and a test that cannot run is
|
|
# not evidence either way.
|
|
- name: Test
|
|
run: cargo test --workspace
|
|
|
|
- name: Build
|
|
run: cargo build --workspace --release
|
|
|
|
android:
|
|
runs-on: linux/amd64
|
|
name: Android (aarch64)
|
|
# Waits for the image build. Without this the pull races the push and the
|
|
# job dies with "manifest unknown" before its first step, which is the
|
|
# failure mode this ordering exists to remove.
|
|
needs: android-image
|
|
container:
|
|
image: gitea.tourolle.paris/dtourolle/darkroom-android:latest
|
|
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@v4
|
|
|
|
# The model, which is in LFS and is not optional.
|
|
#
|
|
# `core/dr-segment/models/*.onnx` is tracked in LFS (.gitattributes), so a
|
|
# plain checkout writes a ~130-byte pointer where 11 MB should be, and
|
|
# `dr-segment`'s build script panics by design rather than embedding a
|
|
# pointer and failing at inference. That failure reads like a broken build
|
|
# instead of a missing fetch, which is how it went unnoticed.
|
|
#
|
|
# Not `lfs: true` on the checkout above: that makes the *checkout* fail on
|
|
# this server — `git lfs fetch` is rejected at
|
|
# `/info/lfs/objects/<oid>` with a client error, because the credential
|
|
# checkout installs for git is not one the LFS endpoint accepts. Fetching
|
|
# it as its own step with an explicit token keeps a credential problem
|
|
# from looking like a checkout problem, and lets the rest of the job say
|
|
# what it thinks.
|
|
#
|
|
# `continue-on-error` deliberately: if this cannot authenticate, the build
|
|
# below still runs and fails with the build script's own message, which
|
|
# names the real problem. A checkout that dies here says nothing.
|
|
- name: Fetch the segmentation model
|
|
continue-on-error: true
|
|
run: |
|
|
set -x
|
|
git lfs install --local
|
|
git config --local lfs.https://gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs.access basic
|
|
git config --local http.extraheader "Authorization: token ${{ secrets.GITEA_TOKEN || github.token }}"
|
|
git lfs pull
|
|
ls -l core/dr-segment/models/
|
|
|
|
- name: Cache cargo
|
|
uses: actions/cache@v4
|
|
with:
|
|
path: |
|
|
/opt/cargo/registry
|
|
target-android
|
|
key: android-${{ hashFiles('**/Cargo.lock') }}
|
|
|
|
# A fast gate on the crates most likely to break the cross-compile, run
|
|
# before the expensive part. It is `cargo check`, so it type-checks
|
|
# without linking and returns in a fraction of the time the step below
|
|
# takes.
|
|
#
|
|
# Not a statement that only these crates cross-compile — `darkroom-android`
|
|
# and the whole UI stack beneath it build for aarch64 too, which is what
|
|
# the API-level step below does. This one exists to fail fast and name a
|
|
# smaller suspect when it does.
|
|
- name: Cross-compile core
|
|
env:
|
|
CARGO_TARGET_DIR: target-android
|
|
run: cargo check -p dr-types -p dr-gpu -p dr-sync --target aarch64-linux-android
|
|
|
|
# The linker targets MIN_API, not the compile SDK. cargo-ndk otherwise
|
|
# defaults to API 21, far below the Vulkan floor this app needs — and the
|
|
# mismatch is invisible until a device refuses to install.
|
|
#
|
|
# Look under the target triple, and fail on a mismatch. Searching the
|
|
# whole target dir for the first `*.so` found the host proc-macro
|
|
# libraries in target-android/debug/deps instead — x86-64 objects built
|
|
# by the runner's gcc, whose .comment section says nothing about Android
|
|
# and can never contradict the expected API. The step passed regardless
|
|
# of what the linker actually did, which is the one thing it exists to
|
|
# rule out.
|
|
- name: Verify minimum API level
|
|
env:
|
|
CARGO_TARGET_DIR: target-android
|
|
run: |
|
|
set -e
|
|
# `darkroom-android`, not a core crate: this step reads the API level
|
|
# out of a *linked* object, and only that crate produces one. It is
|
|
# the workspace's single `crate-type = ["cdylib"]`; a library crate
|
|
# builds an rlib, which is an archive of object files that no linker
|
|
# has yet touched and that `file` therefore has nothing to say about.
|
|
# Asking for `-p dr-gpu` here could only ever reach the "no aarch64
|
|
# .so was produced" branch below, whatever the linker did.
|
|
#
|
|
# It is also the honest artefact to check: the .so this names is the
|
|
# one that ships in the APK, so the API level verified here is the
|
|
# API level a device will refuse to install against.
|
|
cargo ndk -t arm64-v8a build -p darkroom-android --release
|
|
MIN_API=$(sed -n 's/^ARG MIN_API=\([0-9]*\).*/\1/p' docker/android/Dockerfile)
|
|
# Empty on both sides would compare equal and pass, so neither side
|
|
# is allowed to be the result of a failed parse.
|
|
if [ -z "$MIN_API" ]; then
|
|
echo "no ARG MIN_API= in docker/android/Dockerfile"
|
|
exit 1
|
|
fi
|
|
SO=$(find target-android/aarch64-linux-android/release -maxdepth 1 -name '*.so' | head -1)
|
|
if [ -z "$SO" ]; then
|
|
echo "no aarch64 .so was produced"
|
|
exit 1
|
|
fi
|
|
echo "checking $SO"
|
|
file "$SO"
|
|
API=$(file "$SO" | sed -n 's/.*for Android \([0-9]*\).*/\1/p')
|
|
if [ -z "$API" ] || [ "$API" != "$MIN_API" ]; then
|
|
echo "FAIL: linked for Android '${API:-unknown}', expected $MIN_API"
|
|
exit 1
|
|
fi
|
|
|
|
layering:
|
|
runs-on: linux/amd64
|
|
name: Layer separation
|
|
# Node for the JS actions, as above. cargo comes from rustup below.
|
|
container:
|
|
image: catthehacker/ubuntu:act-latest
|
|
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@v4
|
|
|
|
# `cargo tree` resolves the dependency graph, so it needs the registry
|
|
# index but no system libraries — this job builds nothing.
|
|
- name: Install Rust 1.92.0
|
|
run: |
|
|
set -e
|
|
curl -fsSL https://sh.rustup.rs | sh -s -- \
|
|
-y --no-modify-path --profile minimal --default-toolchain 1.92.0
|
|
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
|
|
|
|
# ARCH §6.5a: no core/ crate may depend on the UI toolkit. One stray
|
|
# `use slint::` costs headless golden-image testing and the
|
|
# one-operation-two-presentations property together, and nothing else
|
|
# would notice.
|
|
- name: Core crates must not depend on the UI
|
|
run: |
|
|
set -e
|
|
FAILED=0
|
|
for crate in dr-types dr-gpu dr-sync; do
|
|
if cargo tree -p "$crate" -e normal 2>/dev/null | grep -qE '\bslint\b|\bi-slint'; then
|
|
echo "FAIL: $crate depends on Slint (ARCH §6.5a)"
|
|
FAILED=1
|
|
else
|
|
echo "ok: $crate"
|
|
fi
|
|
done
|
|
exit $FAILED
|