Skip to content

ci: use CPU torch on GitHub runners to stop random CUDA-preload segfault - #807

Merged
Hananel-Hazan merged 1 commit into
masterfrom
ci/cpu-torch
Oct 2, 2026
Merged

Hananel-Hazan merged 1 commit into
masterfrom
ci/cpu-torch

Conversation

@Hananel-Hazan

Copy link
Copy Markdown
Collaborator

Problem

Tests on GitHub runners crash at random with Fatal Python error: Segmentation fault inside torch._preload_cuda_lib (torch/__init__.py:348), before any test runs. The runners have no GPU, but poetry.lock installs the CUDA 13.0 build of torch. Seen on master 7cac0c3 and on Dependabot PRs #804 and #805 (lock-only bumps of tornado/virtualenv, unrelated to torch).

Fix

.github/scripts/use_cpu_torch.sh runs after poetry install in both test workflows. It reads the installed torch/torchvision versions from package metadata (without importing torch), reinstalls those same versions from the CPU wheel index, and fails the job if the resulting torch still reports CUDA.

Local installs and poetry.lock are unchanged.

🤖 Generated with Claude Code

The cu130 torch build pinned in poetry.lock sometimes segfaults in
torch._preload_cuda_lib on GPU-less runners (seen on master 7cac0c3 and
Dependabot PRs #804/#805). After poetry install, reinstall the same
torch/torchvision versions from the CPU wheel index and assert the
build has no CUDA.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Hananel-Hazan
Hananel-Hazan merged commit 247ab83 into master Oct 2, 2026
10 checks passed
Hananel-Hazan added a commit that referenced this pull request Oct 2, 2026
Conflicts came from #807. pythonpackage.yml stays deleted (this branch
replaces it); both jobs in python-app.yml now run
.github/scripts/use_cpu_torch.sh after poetry install.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant