Skip to content

feat: update to 0.19.1 - #287

Merged
velaraptor-runpod merged 8 commits into
mainfrom
feat/0.19.1
May 1, 2026
Merged

velaraptor-runpod merged 8 commits into
mainfrom
feat/0.19.1

Conversation

@velaraptor-runpod

Copy link
Copy Markdown
Contributor

No description provided.

velaraptor-runpod and others added 3 commits April 30, 2026 17:41
- Bump vllm[flashinfer] to 0.19.1 in Dockerfile
- Add OpenAIServingRender (new required dependency in 0.19.x serving layer)
- Pass openai_serving_render to all four serving class constructors
- Remove log_error_stack param (removed upstream in 0.19.x)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…zero adapters

Fixes FDE-194. Previously a malformed LORA_MODULES value was swallowed at
info level and the engine would start with no LoRA adapters, causing 500s
on any request using an adapter model name (e.g. npc-sim-*).

Changes:
- Log at error level when LORA_MODULES cannot be parsed as JSON
- Log at error level when individual adapter dicts fail LoRAModulePath validation
- Log a final error when all adapters fail to load so the cause is obvious
- Accept a single adapter dict (not just an array) for convenience
- Return early when LORA_MODULES is unset to skip unnecessary parsing

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…asing

Fixes FDE-174. Some model stores (e.g. RunPod pre-cached network volumes)
normalize repo IDs to lowercase. HuggingFace Hub caches using the original
casing, so MODEL_NAME=Qwen/Qwen2.5-Coder-32B-Instruct-AWQ would miss a
cache stored as models--qwen--qwen2.5-coder-32b-instruct-awq/ and attempt
a redundant download that fails on limited container storage.

If the exact-case HF cache directory is absent but a lowercase variant
exists, the latest snapshot path is returned directly so vLLM loads from
disk. Absolute paths and models with no lowercase cache are unchanged.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@velaraptor-runpod

Copy link
Copy Markdown
Contributor Author

ok tested.
This env variable needs to be set since KV Cache is done differently. Going to add to hub default settings and READMEs.

PYTORCH_ALLOC_CONF=expandable_segments:True

@velaraptor-runpod
velaraptor-runpod requested a review from jhcipar May 1, 2026 19:14
@velaraptor-runpod

Copy link
Copy Markdown
Contributor Author

@jhcipar

@velaraptor-runpod
velaraptor-runpod merged commit 8a099c1 into main May 1, 2026
4 checks passed
@velaraptor-runpod
velaraptor-runpod deleted the feat/0.19.1 branch May 1, 2026 19:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants