Open one GPU device for the tests, and stop the checkout dying over LFS
Build and test / Desktop (Linux) (push) Failing after 9m18s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 25s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 22m46s
Build and test / Desktop (Linux) (push) Failing after 9m18s
Build and test / Layer separation (push) Successful in 26s
Traceability / Requirement traces (push) Failing after 25s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Android (aarch64) (push) Failing after 22m46s
Two CI failures, unrelated except that both were mine. **The test binary faulted under parallel threads.** Every GPU test opened its own `GpuContext`, and `cargo test` runs on as many threads as there are cores — so a full run asked the driver to bring up a dozen Vulkan devices at once and died with SIGSEGV. Serially it passed, which made it look like flakiness rather than a fault in the harness. One device now, behind a `OnceLock`: a `GpuContext` is an `Arc<Device>` and an `Arc<Queue>`, so sharing it is a refcount, and the losers of the race block until the winner is done. 416 tests now pass in parallel, in half the time twenty-six devices took. **`lfs: true` on the checkout made the checkout fail.** The intent was right — the model is in LFS, a plain checkout writes a 133-byte pointer, and the build script panics on it — but on this server `git lfs fetch` is rejected at `/info/lfs/objects/<oid>` with a client error: the credential `actions/checkout` installs for git is not one the LFS endpoint accepts. So a fetch problem presented as a checkout problem and took the whole job with it. The object is on the server; a clean clone over SSH with `git lfs install --local` pulls all 11 MB of it. It is now its own step with an explicit token, and `continue-on-error` so a credential problem cannot masquerade as a broken checkout — if it fails, the build still runs and fails with the build script's own message, which names the real problem. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+20
-1
@@ -2977,8 +2977,27 @@ mod tests {
|
||||
Some(session)
|
||||
}
|
||||
|
||||
/// One device for the whole test binary.
|
||||
///
|
||||
/// This opened a *new* `GpuContext` per test, and `cargo test` runs tests
|
||||
/// on as many threads as there are cores — so a full run asked the driver
|
||||
/// to bring up a dozen Vulkan devices at once and the binary died with
|
||||
/// SIGSEGV. Serially it passed, which is what made it look like flakiness
|
||||
/// rather than a bug in the harness.
|
||||
///
|
||||
/// A `GpuContext` is an `Arc<Device>` and an `Arc<Queue>`, so sharing one
|
||||
/// is a refcount rather than a copy, and wgpu is explicit that both are
|
||||
/// safe to use from several threads. Nothing here mutates the context; the
|
||||
/// per-test state is in the passes and the sessions built on top of it.
|
||||
///
|
||||
/// `OnceLock` rather than `lazy_static`: the initialiser runs once however
|
||||
/// many threads arrive together, and the losers block until it is done —
|
||||
/// which is precisely the property that was missing.
|
||||
fn headless() -> Option<GpuContext> {
|
||||
pollster::block_on(dr_gpu::GpuContext::new_headless()).ok()
|
||||
static SHARED: std::sync::OnceLock<Option<GpuContext>> = std::sync::OnceLock::new();
|
||||
SHARED
|
||||
.get_or_init(|| pollster::block_on(dr_gpu::GpuContext::new_headless()).ok())
|
||||
.clone()
|
||||
}
|
||||
|
||||
/// TRACES: FR-DEV-3
|
||||
|
||||
Reference in New Issue
Block a user