Files
DarkRoom/.gitea/workflows/build-and-test.yml
T
dtourolleandClaude Opus 5 33847a0bbc
Build and test / Desktop (Linux) (push) Failing after 9m18s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 25s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 22m46s
Open one GPU device for the tests, and stop the checkout dying over LFS
Two CI failures, unrelated except that both were mine.

**The test binary faulted under parallel threads.** Every GPU test opened its
own `GpuContext`, and `cargo test` runs on as many threads as there are cores —
so a full run asked the driver to bring up a dozen Vulkan devices at once and
died with SIGSEGV. Serially it passed, which made it look like flakiness rather
than a fault in the harness. One device now, behind a `OnceLock`: a
`GpuContext` is an `Arc<Device>` and an `Arc<Queue>`, so sharing it is a
refcount, and the losers of the race block until the winner is done. 416 tests
now pass in parallel, in half the time twenty-six devices took.

**`lfs: true` on the checkout made the checkout fail.** The intent was right —
the model is in LFS, a plain checkout writes a 133-byte pointer, and the build
script panics on it — but on this server `git lfs fetch` is rejected at
`/info/lfs/objects/<oid>` with a client error: the credential `actions/checkout`
installs for git is not one the LFS endpoint accepts. So a fetch problem
presented as a checkout problem and took the whole job with it.

The object is on the server; a clean clone over SSH with `git lfs install
--local` pulls all 11 MB of it. It is now its own step with an explicit token,
and `continue-on-error` so a credential problem cannot masquerade as a broken
checkout — if it fails, the build still runs and fails with the build script's
own message, which names the real problem.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 19:36:34 +02:00

254 lines
11 KiB
YAML

name: Build and test
# Desktop and Android are built on every push, per the v0.1 decision to carry
# both platforms from the first commit. An Android break is then caught the day
# it lands rather than at a porting milestone.
on:
push:
branches: [main, master, develop]
pull_request:
branches: [main, master, develop]
jobs:
# The Android job runs inside an image that this repo builds. Ensure it is in
# the registry before anything tries to pull it — see android-image.yml for
# why this is a job rather than a documented manual step. It is a no-op of a
# few seconds unless docker/android actually changed.
android-image:
uses: ./.gitea/workflows/android-image.yml
desktop:
runs-on: linux/amd64
name: Desktop (Linux)
# actions/checkout and actions/cache are JavaScript actions: the runner
# executes them with Node from inside this container. The bare runner image
# has none, so the job failed at checkout before reaching any build step.
container:
image: catthehacker/ubuntu:act-latest
steps:
- name: Checkout
uses: actions/checkout@v4
# The model, which is in LFS and is not optional.
#
# `core/dr-segment/models/*.onnx` is tracked in LFS (.gitattributes), so a
# plain checkout writes a ~130-byte pointer where 11 MB should be, and
# `dr-segment`'s build script panics by design rather than embedding a
# pointer and failing at inference. That failure reads like a broken build
# instead of a missing fetch, which is how it went unnoticed.
#
# Not `lfs: true` on the checkout above: that makes the *checkout* fail on
# this server — `git lfs fetch` is rejected at
# `/info/lfs/objects/<oid>` with a client error, because the credential
# checkout installs for git is not one the LFS endpoint accepts. Fetching
# it as its own step with an explicit token keeps a credential problem
# from looking like a checkout problem, and lets the rest of the job say
# what it thinks.
#
# `continue-on-error` deliberately: if this cannot authenticate, the build
# below still runs and fails with the build script's own message, which
# names the real problem. A checkout that dies here says nothing.
- name: Fetch the segmentation model
continue-on-error: true
run: |
set -x
git lfs install --local
git config --local lfs.https://gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs.access basic
git config --local http.extraheader "Authorization: token ${{ secrets.GITEA_TOKEN || github.token }}"
git lfs pull
ls -l core/dr-segment/models/
- name: Cache cargo
uses: actions/cache@v4
with:
path: |
~/.cargo/registry
~/.cargo/git
target
key: desktop-${{ runner.os }}-${{ hashFiles('**/Cargo.lock') }}
# Slint and winit need these at build time; the runner image is minimal.
- name: Build dependencies
run: |
apt-get update -qq
apt-get install -y -qq pkg-config libfontconfig1-dev libxkbcommon-dev
# The act image ships Node but no Rust. Pinned to the workspace
# rust-version so CI, the Android image, and local builds agree — a
# floating toolchain turns an unrelated push into a mystery failure.
- name: Install Rust 1.92.0
run: |
set -e
curl -fsSL https://sh.rustup.rs | sh -s -- \
-y --no-modify-path --profile minimal \
--default-toolchain 1.92.0 --component rustfmt,clippy
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
- name: Format
run: cargo fmt --all -- --check
- name: Clippy
run: cargo clippy --workspace --all-targets -- -D warnings
# GPU tests skip themselves where no adapter is present rather than
# failing — CI runners generally have none, and a test that cannot run is
# not evidence either way.
- name: Test
run: cargo test --workspace
- name: Build
run: cargo build --workspace --release
android:
runs-on: linux/amd64
name: Android (aarch64)
# Waits for the image build. Without this the pull races the push and the
# job dies with "manifest unknown" before its first step, which is the
# failure mode this ordering exists to remove.
needs: android-image
container:
image: gitea.tourolle.paris/dtourolle/darkroom-android:latest
steps:
- name: Checkout
uses: actions/checkout@v4
# The model, which is in LFS and is not optional.
#
# `core/dr-segment/models/*.onnx` is tracked in LFS (.gitattributes), so a
# plain checkout writes a ~130-byte pointer where 11 MB should be, and
# `dr-segment`'s build script panics by design rather than embedding a
# pointer and failing at inference. That failure reads like a broken build
# instead of a missing fetch, which is how it went unnoticed.
#
# Not `lfs: true` on the checkout above: that makes the *checkout* fail on
# this server — `git lfs fetch` is rejected at
# `/info/lfs/objects/<oid>` with a client error, because the credential
# checkout installs for git is not one the LFS endpoint accepts. Fetching
# it as its own step with an explicit token keeps a credential problem
# from looking like a checkout problem, and lets the rest of the job say
# what it thinks.
#
# `continue-on-error` deliberately: if this cannot authenticate, the build
# below still runs and fails with the build script's own message, which
# names the real problem. A checkout that dies here says nothing.
- name: Fetch the segmentation model
continue-on-error: true
run: |
set -x
git lfs install --local
git config --local lfs.https://gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs.access basic
git config --local http.extraheader "Authorization: token ${{ secrets.GITEA_TOKEN || github.token }}"
git lfs pull
ls -l core/dr-segment/models/
- name: Cache cargo
uses: actions/cache@v4
with:
path: |
/opt/cargo/registry
target-android
key: android-${{ hashFiles('**/Cargo.lock') }}
# A fast gate on the crates most likely to break the cross-compile, run
# before the expensive part. It is `cargo check`, so it type-checks
# without linking and returns in a fraction of the time the step below
# takes.
#
# Not a statement that only these crates cross-compile — `darkroom-android`
# and the whole UI stack beneath it build for aarch64 too, which is what
# the API-level step below does. This one exists to fail fast and name a
# smaller suspect when it does.
- name: Cross-compile core
env:
CARGO_TARGET_DIR: target-android
run: cargo check -p dr-types -p dr-gpu -p dr-sync --target aarch64-linux-android
# The linker targets MIN_API, not the compile SDK. cargo-ndk otherwise
# defaults to API 21, far below the Vulkan floor this app needs — and the
# mismatch is invisible until a device refuses to install.
#
# Look under the target triple, and fail on a mismatch. Searching the
# whole target dir for the first `*.so` found the host proc-macro
# libraries in target-android/debug/deps instead — x86-64 objects built
# by the runner's gcc, whose .comment section says nothing about Android
# and can never contradict the expected API. The step passed regardless
# of what the linker actually did, which is the one thing it exists to
# rule out.
- name: Verify minimum API level
env:
CARGO_TARGET_DIR: target-android
run: |
set -e
# `darkroom-android`, not a core crate: this step reads the API level
# out of a *linked* object, and only that crate produces one. It is
# the workspace's single `crate-type = ["cdylib"]`; a library crate
# builds an rlib, which is an archive of object files that no linker
# has yet touched and that `file` therefore has nothing to say about.
# Asking for `-p dr-gpu` here could only ever reach the "no aarch64
# .so was produced" branch below, whatever the linker did.
#
# It is also the honest artefact to check: the .so this names is the
# one that ships in the APK, so the API level verified here is the
# API level a device will refuse to install against.
cargo ndk -t arm64-v8a build -p darkroom-android --release
MIN_API=$(sed -n 's/^ARG MIN_API=\([0-9]*\).*/\1/p' docker/android/Dockerfile)
# Empty on both sides would compare equal and pass, so neither side
# is allowed to be the result of a failed parse.
if [ -z "$MIN_API" ]; then
echo "no ARG MIN_API= in docker/android/Dockerfile"
exit 1
fi
SO=$(find target-android/aarch64-linux-android/release -maxdepth 1 -name '*.so' | head -1)
if [ -z "$SO" ]; then
echo "no aarch64 .so was produced"
exit 1
fi
echo "checking $SO"
file "$SO"
API=$(file "$SO" | sed -n 's/.*for Android \([0-9]*\).*/\1/p')
if [ -z "$API" ] || [ "$API" != "$MIN_API" ]; then
echo "FAIL: linked for Android '${API:-unknown}', expected $MIN_API"
exit 1
fi
layering:
runs-on: linux/amd64
name: Layer separation
# Node for the JS actions, as above. cargo comes from rustup below.
container:
image: catthehacker/ubuntu:act-latest
steps:
- name: Checkout
uses: actions/checkout@v4
# `cargo tree` resolves the dependency graph, so it needs the registry
# index but no system libraries — this job builds nothing.
- name: Install Rust 1.92.0
run: |
set -e
curl -fsSL https://sh.rustup.rs | sh -s -- \
-y --no-modify-path --profile minimal --default-toolchain 1.92.0
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
# ARCH §6.5a: no core/ crate may depend on the UI toolkit. One stray
# `use slint::` costs headless golden-image testing and the
# one-operation-two-presentations property together, and nothing else
# would notice.
- name: Core crates must not depend on the UI
run: |
set -e
FAILED=0
for crate in dr-types dr-gpu dr-sync; do
if cargo tree -p "$crate" -e normal 2>/dev/null | grep -qE '\bslint\b|\bi-slint'; then
echo "FAIL: $crate depends on Slint (ARCH §6.5a)"
FAILED=1
else
echo "ok: $crate"
fi
done
exit $FAILED