Skip to content

fix: make .runpod/tests.json hub tests pass on CUDA 13.0 - #299

Merged
TimPietruskyRunPod merged 1 commit into
mainfrom
fix/hub-tests-cuda13-kernels-timeout
Jun 2, 2026
Merged

TimPietruskyRunPod merged 1 commit into
mainfrom
fix/hub-tests-cuda13-kernels-timeout

Conversation

@TimPietruskyRunPod

Copy link
Copy Markdown
Contributor

Summary

Three coupled fixes so the re-enabled hub tests from #295 actually pass. Verified end-to-end on a private fork — both basic_inference_test and openai_messages_test green on v0.1.3.

What was broken

After v2.20.0 (#288) bumped the base image to CUDA 13.0, the hub-test flow had three independent failures stacked on top of each other. Each one masked the next:

  1. CUDA version mismatch in tests.json — allowedCudaVersions was still ["12.9", "12.8", "12.7", "12.6", "12.5"] so the test pod was scheduled on a driver < 13.0 and container init aborted at the nvidia-container-cli hook with unsatisfied condition: cuda>=13.0.

  2. kernels 0.15.1 breaking transformers — kernels v0.15.0 (#544) tightened LayerRepository to require revision= or version=. transformers>=5 still constructs LayerRepository(repo_id=..., layer_name=...) without either (hub_kernels.py:89), so worker import raises ValueError at module-load time. Pinning kernels<0.15 keeps us on 0.14.1 which doesn't have the strict check.

  3. Test timeout too tight for vLLM cold start — vLLM 0.20.2 torch.compile + FlashInfer warmup on RTX 4090 takes ~60–70s before the first inference is served (observed: torch.compile took 19.34 s + warmup + model load). The 30s per-test timeout fired before the worker came up, producing context cancelled or timed out: context deadline exceeded for every test even when the worker was healthy. Bumping to 300s.

Changes

  • .runpod/tests.json: allowedCudaVersions → ["13.0"]; both test timeout → 300000
  • builder/requirements.txt: kernels → kernels<0.15

Validation

Private fork TimPietruskyRunPod/worker-vllm, identical code path, four iterations to isolate each bug:

Tag Change Result
v0.1.0 upstream mirror Container init fail — cuda>=13.0
v0.1.1 + CUDA 13.0 Worker boot fail — kernels.LayerRepository ValueError
v0.1.2 + kernels<0.15 vLLM up (torch.compile 19.34s), tests timeout @ 30s
v0.1.3 + 300s timeout 2/2 tests passing

Refs: DR-1161

Test plan

  • PR build kicks off (dev-refs-pull-<N>-merge)
  • Hub tests pass on this PR's image: basic_inference_test ✓ and openai_messages_test ✓
  • After merge + version bump release, the same tests pass on runpod/worker-v1-vllm:vX.Y.Z

Three coupled fixes verified end-to-end on a private fork
(TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing):

1. tests.json allowedCudaVersions: 12.x → 13.0
   The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0
   (#288, #289), but tests.json was still pinned to 12.5–12.9, so
   the test pod was scheduled on a GPU with driver < 13.0 and
   container init failed at the nvidia-container-cli hook with
   "unsatisfied condition: cuda>=13.0".

2. requirements.txt kernels<0.15
   huggingface/kernels v0.15.1 tightened LayerRepository to require
   a revision or version argument
   (huggingface/kernels#544). transformers
   >=5 still constructs LayerRepository(repo_id=..., layer_name=...)
   without either, so worker import raised ValueError during
   `from transformers import ...`. 0.14.1 is the last safe release.

3. tests.json timeout 30000 → 300000
   vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090
   for SmolLM2-135M takes ~60–70s before the first request can be
   served. The previous 30s per-test timeout fired before the
   worker came up, producing "context cancelled or timed out:
   context deadline exceeded" for every test even when the worker
   was healthy. 300s gives enough headroom for cold start + the
   actual inference call.

Refs: DR-1161
@TimPietruskyRunPod
TimPietruskyRunPod merged commit dac05b6 into main Jun 2, 2026
5 checks passed
@TimPietruskyRunPod
TimPietruskyRunPod deleted the fix/hub-tests-cuda13-kernels-timeout branch June 2, 2026 15:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants