Repository navigation
fix: make .runpod/tests.json hub tests pass on CUDA 13.0 - #299
Merged
Merged
Conversation
Three coupled fixes verified end-to-end on a private fork (TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing): 1. tests.json allowedCudaVersions: 12.x → 13.0 The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0 (#288, #289), but tests.json was still pinned to 12.5–12.9, so the test pod was scheduled on a GPU with driver < 13.0 and container init failed at the nvidia-container-cli hook with "unsatisfied condition: cuda>=13.0". 2. requirements.txt kernels<0.15 huggingface/kernels v0.15.1 tightened LayerRepository to require a revision or version argument (huggingface/kernels#544). transformers >=5 still constructs LayerRepository(repo_id=..., layer_name=...) without either, so worker import raised ValueError during `from transformers import ...`. 0.14.1 is the last safe release. 3. tests.json timeout 30000 → 300000 vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090 for SmolLM2-135M takes ~60–70s before the first request can be served. The previous 30s per-test timeout fired before the worker came up, producing "context cancelled or timed out: context deadline exceeded" for every test even when the worker was healthy. 300s gives enough headroom for cold start + the actual inference call. Refs: DR-1161
velaraptor-runpod
approved these changes
Jun 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Three coupled fixes so the re-enabled hub tests from #295 actually pass. Verified end-to-end on a private fork — both
basic_inference_testandopenai_messages_testgreen on v0.1.3.What was broken
After v2.20.0 (#288) bumped the base image to CUDA 13.0, the hub-test flow had three independent failures stacked on top of each other. Each one masked the next:
CUDA version mismatch in
tests.json—allowedCudaVersionswas still["12.9", "12.8", "12.7", "12.6", "12.5"]so the test pod was scheduled on a driver < 13.0 and container init aborted at thenvidia-container-clihook withunsatisfied condition: cuda>=13.0.kernels0.15.1 breakingtransformers—kernelsv0.15.0 (#544) tightenedLayerRepositoryto requirerevision=orversion=.transformers>=5still constructsLayerRepository(repo_id=..., layer_name=...)without either (hub_kernels.py:89), so worker import raisesValueErrorat module-load time. Pinningkernels<0.15keeps us on 0.14.1 which doesn't have the strict check.Test timeout too tight for vLLM cold start — vLLM 0.20.2
torch.compile+ FlashInfer warmup on RTX 4090 takes ~60–70s before the first inference is served (observed:torch.compile took 19.34 s+ warmup + model load). The 30s per-test timeout fired before the worker came up, producingcontext cancelled or timed out: context deadline exceededfor every test even when the worker was healthy. Bumping to 300s.Changes
.runpod/tests.json:allowedCudaVersions→["13.0"]; both testtimeout→300000builder/requirements.txt:kernels→kernels<0.15Validation
Private fork
TimPietruskyRunPod/worker-vllm, identical code path, four iterations to isolate each bug:cuda>=13.0kernels.LayerRepositoryValueErrorkernels<0.15torch.compile 19.34s), tests timeout @ 30sRefs: DR-1161
Test plan
dev-refs-pull-<N>-merge)basic_inference_test✓ andopenai_messages_test✓runpod/worker-v1-vllm:vX.Y.Z