diff --git a/.runpod/README.md b/.runpod/README.md index 232dd701..460a7e51 100644 --- a/.runpod/README.md +++ b/.runpod/README.md @@ -6,7 +6,7 @@ Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API [![RunPod](https://api.runpod.io/badge/runpod-workers/worker-vllm)](https://www.runpod.io/console/hub/runpod-workers/worker-vllm) -Current vLLM version: [0.29.0](https://github.com/vllm-project/vllm/releases/tag/v0.29.0) +Current vLLM version: [0.28.0](https://github.com/vllm-project/vllm/releases/tag/v0.28.0) --- diff --git a/Dockerfile b/Dockerfile index b0253baf..a7b916a1 100644 --- a/Dockerfile +++ b/Dockerfile @@ -1,7 +1,7 @@ # Worker image = official vLLM OpenAI server image + RunPod serverless wrapper. # vLLM upgrades are now a single build ARG: # docker buildx build --build-arg VLLM_VERSION=v0.23.0 ... -ARG VLLM_VERSION=v0.29.0 +ARG VLLM_VERSION=v0.28.0 FROM vllm/vllm-openai:${VLLM_VERSION} # Re-declare so the stage can reference it in RUN steps below. ARG VLLM_VERSION diff --git a/README.md b/README.md index 75ca1a06..50dc0800 100644 --- a/README.md +++ b/README.md @@ -8,7 +8,7 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https: ![vLLM worker banner](https://image.runpod.ai/preview/vllm/vllm-banner.png) -Current vLLM version: [0.29.0](https://github.com/vllm-project/vllm/releases/tag/v0.29.0) +Current vLLM version: [0.28.0](https://github.com/vllm-project/vllm/releases/tag/v0.28.0) > Want a **load balancing** endpoint (direct HTTP, no job queue)? You don't need this worker — deploy the official vLLM image as-is. See [Option 3: Load Balancing with the vLLM Image](#option-3-load-balancing-with-the-vllm-image).