Open one GPU device for the tests, and stop the checkout dying over LFS
Build and test / Desktop (Linux) (push) Failing after 9m18s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 25s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 22m46s
Build and test / Desktop (Linux) (push) Failing after 9m18s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 25s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 22m46s
Two CI failures, unrelated except that both were mine. **The test binary faulted under parallel threads.** Every GPU test opened its own `GpuContext`, and `cargo test` runs on as many threads as there are cores — so a full run asked the driver to bring up a dozen Vulkan devices at once and died with SIGSEGV. Serially it passed, which made it look like flakiness rather than a fault in the harness. One device now, behind a `OnceLock`: a `GpuContext` is an `Arc<Device>` and an `Arc<Queue>`, so sharing it is a refcount, and the losers of the race block until the winner is done. 416 tests now pass in parallel, in half the time twenty-six devices took. **`lfs: true` on the checkout made the checkout fail.** The intent was right — the model is in LFS, a plain checkout writes a 133-byte pointer, and the build script panics on it — but on this server `git lfs fetch` is rejected at `/info/lfs/objects/<oid>` with a client error: the credential `actions/checkout` installs for git is not one the LFS endpoint accepts. So a fetch problem presented as a checkout problem and took the whole job with it. The object is on the server; a clean clone over SSH with `git lfs install --local` pulls all 11 MB of it. It is now its own step with an explicit token, and `continue-on-error` so a credential problem cannot masquerade as a broken checkout — if it fails, the build still runs and fails with the build script's own message, which names the real problem. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -29,18 +29,36 @@ jobs:
|
||||
|
||||
steps:
|
||||
- name: Checkout
|
||||
# `lfs: true` is load-bearing, not hygiene. `core/dr-segment/models/*.onnx`
|
||||
# is tracked in LFS (see .gitattributes), and a checkout without it writes
|
||||
# a ~130-byte pointer where the 11 MB model should be. `dr-segment`'s build
|
||||
# script detects that and panics by design, so every job that compiles it
|
||||
# fails at the build script with a message about git-lfs — which reads like
|
||||
# a broken build rather than a missing fetch.
|
||||
#
|
||||
# It stayed hidden while the Android job built `-p dr-gpu`, which does not
|
||||
# reach dr-segment. Building the app crate does.
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
lfs: true
|
||||
|
||||
# The model, which is in LFS and is not optional.
|
||||
#
|
||||
# `core/dr-segment/models/*.onnx` is tracked in LFS (.gitattributes), so a
|
||||
# plain checkout writes a ~130-byte pointer where 11 MB should be, and
|
||||
# `dr-segment`'s build script panics by design rather than embedding a
|
||||
# pointer and failing at inference. That failure reads like a broken build
|
||||
# instead of a missing fetch, which is how it went unnoticed.
|
||||
#
|
||||
# Not `lfs: true` on the checkout above: that makes the *checkout* fail on
|
||||
# this server — `git lfs fetch` is rejected at
|
||||
# `/info/lfs/objects/<oid>` with a client error, because the credential
|
||||
# checkout installs for git is not one the LFS endpoint accepts. Fetching
|
||||
# it as its own step with an explicit token keeps a credential problem
|
||||
# from looking like a checkout problem, and lets the rest of the job say
|
||||
# what it thinks.
|
||||
#
|
||||
# `continue-on-error` deliberately: if this cannot authenticate, the build
|
||||
# below still runs and fails with the build script's own message, which
|
||||
# names the real problem. A checkout that dies here says nothing.
|
||||
- name: Fetch the segmentation model
|
||||
continue-on-error: true
|
||||
run: |
|
||||
set -x
|
||||
git lfs install --local
|
||||
git config --local lfs.https://gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs.access basic
|
||||
git config --local http.extraheader "Authorization: token ${{ secrets.GITEA_TOKEN || github.token }}"
|
||||
git lfs pull
|
||||
ls -l core/dr-segment/models/
|
||||
|
||||
- name: Cache cargo
|
||||
uses: actions/cache@v4
|
||||
@@ -95,18 +113,36 @@ jobs:
|
||||
|
||||
steps:
|
||||
- name: Checkout
|
||||
# `lfs: true` is load-bearing, not hygiene. `core/dr-segment/models/*.onnx`
|
||||
# is tracked in LFS (see .gitattributes), and a checkout without it writes
|
||||
# a ~130-byte pointer where the 11 MB model should be. `dr-segment`'s build
|
||||
# script detects that and panics by design, so every job that compiles it
|
||||
# fails at the build script with a message about git-lfs — which reads like
|
||||
# a broken build rather than a missing fetch.
|
||||
#
|
||||
# It stayed hidden while the Android job built `-p dr-gpu`, which does not
|
||||
# reach dr-segment. Building the app crate does.
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
lfs: true
|
||||
|
||||
# The model, which is in LFS and is not optional.
|
||||
#
|
||||
# `core/dr-segment/models/*.onnx` is tracked in LFS (.gitattributes), so a
|
||||
# plain checkout writes a ~130-byte pointer where 11 MB should be, and
|
||||
# `dr-segment`'s build script panics by design rather than embedding a
|
||||
# pointer and failing at inference. That failure reads like a broken build
|
||||
# instead of a missing fetch, which is how it went unnoticed.
|
||||
#
|
||||
# Not `lfs: true` on the checkout above: that makes the *checkout* fail on
|
||||
# this server — `git lfs fetch` is rejected at
|
||||
# `/info/lfs/objects/<oid>` with a client error, because the credential
|
||||
# checkout installs for git is not one the LFS endpoint accepts. Fetching
|
||||
# it as its own step with an explicit token keeps a credential problem
|
||||
# from looking like a checkout problem, and lets the rest of the job say
|
||||
# what it thinks.
|
||||
#
|
||||
# `continue-on-error` deliberately: if this cannot authenticate, the build
|
||||
# below still runs and fails with the build script's own message, which
|
||||
# names the real problem. A checkout that dies here says nothing.
|
||||
- name: Fetch the segmentation model
|
||||
continue-on-error: true
|
||||
run: |
|
||||
set -x
|
||||
git lfs install --local
|
||||
git config --local lfs.https://gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs.access basic
|
||||
git config --local http.extraheader "Authorization: token ${{ secrets.GITEA_TOKEN || github.token }}"
|
||||
git lfs pull
|
||||
ls -l core/dr-segment/models/
|
||||
|
||||
- name: Cache cargo
|
||||
uses: actions/cache@v4
|
||||
|
||||
Reference in New Issue
Block a user