diff --git a/.runpod/README.md b/.runpod/README.md index dceb3275..6790449b 100644 --- a/.runpod/README.md +++ b/.runpod/README.md @@ -6,7 +6,7 @@ Run LLMs using [vLLM](https://docs.vllm.ai) with an OpenAI-compatible API [![RunPod](https://api.runpod.io/badge/runpod-workers/worker-vllm)](https://www.runpod.io/console/hub/runpod-workers/worker-vllm) -Current vLLM version: [0.30.0](https://github.com/vllm-project/vllm/releases/tag/v0.30.0) +Current vLLM version: [0.31.0](https://github.com/vllm-project/vllm/releases/tag/v0.31.0) --- diff --git a/README.md b/README.md index cce51bbe..4c51c8f1 100644 --- a/README.md +++ b/README.md @@ -8,7 +8,7 @@ Deploy OpenAI-Compatible Blazing-Fast LLM Endpoints powered by the [vLLM](https: ![vLLM worker banner](https://image.runpod.ai/preview/vllm/vllm-banner.png) -Current vLLM version: [0.30.0](https://github.com/vllm-project/vllm/releases/tag/v0.30.0) +Current vLLM version: [0.31.0](https://github.com/vllm-project/vllm/releases/tag/v0.31.0) > Want a **load balancing** endpoint (direct HTTP, no job queue)? You don't need this worker — deploy the official vLLM image as-is. See [Option 3: Load Balancing with the vLLM Image](#option-3-load-balancing-with-the-vllm-image).