From b0e946394efd5bb0f1a658640a39ea21fd44611a Mon Sep 17 00:00:00 2001 From: Luke Piette Date: Mon, 21 Sep 2026 14:14:31 -0700 Subject: [PATCH] Roll the engine back to vLLM v0.28.0 until the v0.29 first-request regressions are fixed Template v2.27.0 (vLLM v0.29.0) fails ~31% of first requests against ~10-11% on v2.26.0 (v0.28.0). Two engine-side regressions dominate the 50 failed endpoints analysed: - "Revision Not Found" (88%): v0.29.0 bundles huggingface_hub 1.28.0, whose resolved-revision pin leaks across repos (model hash applied to the tokenizer/weights repo -> guaranteed 404); fixed upstream only in vllm#56092 + hf_hub >= 1.30, which no 0.29.x ships. - CUDA OOM (58%): v0.29 costs ~1-2 GiB more init headroom (peak activation +1.13 GiB and an up-front cudagraph reserve, measured A/B on identical A40s with the same model and env), flipping marginal models from "fits" to "OOM at init". Releasing this as v2.27.1 makes the Hub default safe again while the fail-fast work lands separately; re-bump to a vLLM release that carries vllm#56092 with huggingface_hub >= 1.30. Co-Authored-By: Claude Fable 5 --- .runpod/README.md | 2 +- Dockerfile | 2 +- README.md | 2 +- 3 files changed, 3 insertions(+), 3 deletions(-) diff --git a/.runpod/README.md b/.runpod/README.md index 232dd701..460a7e51 100644 --- a/.runpod/README.md +++ b/.runpod/README.md @@ -6,7 +6,7 @@ Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API [![RunPod](https://api.runpod.io/badge/runpod-workers/worker-vllm)](https://www.runpod.io/console/hub/runpod-workers/worker-vllm) -Current vLLM version: [0.29.0](https://github.com/vllm-project/vllm/releases/tag/v0.29.0) +Current vLLM version: [0.28.0](https://github.com/vllm-project/vllm/releases/tag/v0.28.0) --- diff --git a/Dockerfile b/Dockerfile index b0253baf..a7b916a1 100644 --- a/Dockerfile +++ b/Dockerfile @@ -1,7 +1,7 @@ # Worker image = official vLLM OpenAI server image + RunPod serverless wrapper. # vLLM upgrades are now a single build ARG: # docker buildx build --build-arg VLLM_VERSION=v0.23.0 ... -ARG VLLM_VERSION=v0.29.0 +ARG VLLM_VERSION=v0.28.0 FROM vllm/vllm-openai:${VLLM_VERSION} # Re-declare so the stage can reference it in RUN steps below. ARG VLLM_VERSION diff --git a/README.md b/README.md index 75ca1a06..50dc0800 100644 --- a/README.md +++ b/README.md @@ -8,7 +8,7 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https: ![vLLM worker banner](https://image.runpod.ai/preview/vllm/vllm-banner.png) -Current vLLM version: [0.29.0](https://github.com/vllm-project/vllm/releases/tag/v0.29.0) +Current vLLM version: [0.28.0](https://github.com/vllm-project/vllm/releases/tag/v0.28.0) > Want a **load balancing** endpoint (direct HTTP, no job queue)? You don't need this worker — deploy the official vLLM image as-is. See [Option 3: Load Balancing with the vLLM Image](#option-3-load-balancing-with-the-vllm-image).