Skip to content

feat: add notifications for new prs, issues, and new releases of vllm - #289

Merged
velaraptor-runpod merged 2 commits into
mainfrom
feat/add-notifications
May 1, 2026
Merged

velaraptor-runpod merged 2 commits into
mainfrom
feat/add-notifications

Conversation

@velaraptor-runpod

Copy link
Copy Markdown
Contributor

No description provided.

@velaraptor-runpod
velaraptor-runpod requested a review from Yhlong00 May 1, 2026 00:55
Comment thread .github/workflows/slack-pr-issue-notify.yml Fixed
Comment thread .github/workflows/slack-pr-issue-notify.yml Fixed
Comment thread .github/workflows/slack-pr-issue-notify.yml Fixed
Comment thread .github/workflows/slack-vllm-monitor.yml Fixed
@velaraptor-runpod
velaraptor-runpod merged commit cff7b09 into main May 1, 2026
4 checks passed
@velaraptor-runpod
velaraptor-runpod deleted the feat/add-notifications branch May 1, 2026 01:03
TimPietruskyRunPod added a commit that referenced this pull request Jun 2, 2026
Three coupled fixes verified end-to-end on a private fork
(TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing):

1. tests.json allowedCudaVersions: 12.x → 13.0
   The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0
   (#288, #289), but tests.json was still pinned to 12.5–12.9, so
   the test pod was scheduled on a GPU with driver < 13.0 and
   container init failed at the nvidia-container-cli hook with
   "unsatisfied condition: cuda>=13.0".

2. requirements.txt kernels<0.15
   huggingface/kernels v0.15.1 tightened LayerRepository to require
   a revision or version argument
   (huggingface/kernels#544). transformers
   >=5 still constructs LayerRepository(repo_id=..., layer_name=...)
   without either, so worker import raised ValueError during
   `from transformers import ...`. 0.14.1 is the last safe release.

3. tests.json timeout 30000 → 300000
   vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090
   for SmolLM2-135M takes ~60–70s before the first request can be
   served. The previous 30s per-test timeout fired before the
   worker came up, producing "context cancelled or timed out:
   context deadline exceeded" for every test even when the worker
   was healthy. 300s gives enough headroom for cold start + the
   actual inference call.

Refs: DR-1161
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants