Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 5 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,13 +1,13 @@
# vLLM Load Balancer

A FastAPI-based load balancer for serving vLLM models with RunPod integration. Provides OpenAI-compatible APIs with streaming and non-streaming text generation.
A FastAPI-based load balancer for serving vLLM models with Runpod integration. Provides OpenAI-compatible APIs with streaming and non-streaming text generation.

## Prerequisites

Before you begin, make sure you have:

- A RunPod account (sign up at [runpod.io](https://runpod.io))
- RunPod API key (available in your RunPod dashboard)
- A Runpod account (sign up at [runpod.io](https://runpod.io))
- Runpod API key (available in your Runpod dashboard)
- Basic understanding of REST APIs and HTTP requests
- `curl` or a similar tool for testing API endpoints

Expand All @@ -17,7 +17,7 @@ Use the pre-built Docker image: `runpod/vllm-loadbalancer:dev`

## Environment Variables

Configure these environment variables in your RunPod endpoint:
Configure these environment variables in your Runpod endpoint:

| Variable | Required | Description | Default | Example |
|----------|----------|-------------|---------|---------|
Expand All @@ -29,7 +29,7 @@ Configure these environment variables in your RunPod endpoint:
| `GPU_MEMORY_UTILIZATION` | No | GPU memory usage ratio | `0.9` | `0.8` |
| `ENFORCE_EAGER` | No | Disable CUDA graphs | `false` | `true` |

## Deployment on RunPod
## Deployment on Runpod

1. Create a new serverless endpoint
2. Use Docker image: `runpod/vllm-loadbalancer:dev`
Expand Down