Open one GPU device for the tests, and stop the checkout dying over LFS
Build and test / Desktop (Linux) (push) Failing after 9m18s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 25s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 22m46s

Two CI failures, unrelated except that both were mine.

**The test binary faulted under parallel threads.** Every GPU test opened its
own `GpuContext`, and `cargo test` runs on as many threads as there are cores —
so a full run asked the driver to bring up a dozen Vulkan devices at once and
died with SIGSEGV. Serially it passed, which made it look like flakiness rather
than a fault in the harness. One device now, behind a `OnceLock`: a
`GpuContext` is an `Arc<Device>` and an `Arc<Queue>`, so sharing it is a
refcount, and the losers of the race block until the winner is done. 416 tests
now pass in parallel, in half the time twenty-six devices took.

**`lfs: true` on the checkout made the checkout fail.** The intent was right —
the model is in LFS, a plain checkout writes a 133-byte pointer, and the build
script panics on it — but on this server `git lfs fetch` is rejected at
`/info/lfs/objects/<oid>` with a client error: the credential `actions/checkout`
installs for git is not one the LFS endpoint accepts. So a fetch problem
presented as a checkout problem and took the whole job with it.

The object is on the server; a clean clone over SSH with `git lfs install
--local` pulls all 11 MB of it. It is now its own step with an explicit token,
and `continue-on-error` so a credential problem cannot masquerade as a broken
checkout — if it fails, the build still runs and fails with the build script's
own message, which names the real problem.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-24 19:36:34 +02:00
co-authored by Claude Opus 5
parent a3785ba55d
commit 33847a0bbc
3 changed files with 115 additions and 60 deletions
+20 -1
View File
@@ -2977,8 +2977,27 @@ mod tests {
Some(session)
}
/// One device for the whole test binary.
///
/// This opened a *new* `GpuContext` per test, and `cargo test` runs tests
/// on as many threads as there are cores — so a full run asked the driver
/// to bring up a dozen Vulkan devices at once and the binary died with
/// SIGSEGV. Serially it passed, which is what made it look like flakiness
/// rather than a bug in the harness.
///
/// A `GpuContext` is an `Arc<Device>` and an `Arc<Queue>`, so sharing one
/// is a refcount rather than a copy, and wgpu is explicit that both are
/// safe to use from several threads. Nothing here mutates the context; the
/// per-test state is in the passes and the sessions built on top of it.
///
/// `OnceLock` rather than `lazy_static`: the initialiser runs once however
/// many threads arrive together, and the losers block until it is done —
/// which is precisely the property that was missing.
fn headless() -> Option<GpuContext> {
pollster::block_on(dr_gpu::GpuContext::new_headless()).ok()
static SHARED: std::sync::OnceLock<Option<GpuContext>> = std::sync::OnceLock::new();
SHARED
.get_or_init(|| pollster::block_on(dr_gpu::GpuContext::new_headless()).ok())
.clone()
}
/// TRACES: FR-DEV-3