Skip to content

Roll the engine back to vLLM v0.28.0 until the v0.29 first-request regressions are fixed - #344

Merged
lukepiette merged 1 commit into
runpod-workers:mainfrom
lukepiette:rollback-vllm-v0.28.0
Sep 21, 2026
Merged

lukepiette merged 1 commit into
runpod-workers:mainfrom
lukepiette:rollback-vllm-v0.28.0

Conversation

@lukepiette

Copy link
Copy Markdown
Contributor

Why

Hub template v2.27.0 (vLLM v0.29.0) fails ~31% of first requests vs ~10–11% on v2.26.0 (v0.28.0), measured Sep 11–17. The only functional change in v2.27.0 was the engine bump, and both dominant failure modes are engine-side:

  1. "Revision Not Found" (88% of failed endpoints). v0.29.0 ships huggingface_hub 1.28.0, whose resolved-revision pin leaks across repos — the model repo's commit hash gets applied to the tokenizer/weights repo, guaranteeing RevisionNotFoundError. Properly fixed only in vllm#56092 + hf_hub ≥ 1.30, which no 0.29.x carries. (Hash pinning itself is deliberate upstream — the Artifact Pin Decay advisory — so only the leak, not the pinning, is fixable.)
  2. CUDA OOM (58%). Measured A/B on identical A40s, same model (Qwen2.5-Coder-32B-AWQ), same env: v0.29 uses +1.13 GiB peak activation (1.48 → 2.61 GiB) and reserves estimated cudagraph memory up front (its own log: requested 0.95 ≡ effective 0.9247), shrinking available KV cache 5% and flipping marginal models from "fits" to "OOM at init".

What

One-line engine pin v0.29.0 → v0.28.0 plus the README version strings. The v2.27.0-era wrapper improvements (startup-error classification, flag-parser compat) are kept — check_vllm_flags/args_builder support both parsers.

Shipping

Merging this and cutting v2.27.1 makes the Hub default template safe again (validated live: the same model/env that boots on v2.26.0 boots on v0.28.0-based images; see the A/B in #343). The fail-fast hardening (revision relaunch, pre-download fit check, HF_TOKEN input) lands separately in #343 and is version-independent. Re-bump the engine once a vLLM release carries vllm#56092 with hf_hub ≥ 1.30.

Re-measure: after v2.27.1 is the Hub default, compare first-request success for endpoints created on v2.27.1 vs the v2.27.0 cohort over one week; target is back to ~10–11% failure.

🤖 Generated with Claude Code

…gressions are fixed

Template v2.27.0 (vLLM v0.29.0) fails ~31% of first requests against
~10-11% on v2.26.0 (v0.28.0). Two engine-side regressions dominate the 50
failed endpoints analysed:

- "Revision Not Found" (88%): v0.29.0 bundles huggingface_hub 1.28.0, whose
  resolved-revision pin leaks across repos (model hash applied to the
  tokenizer/weights repo -> guaranteed 404); fixed upstream only in
  vllm#56092 + hf_hub >= 1.30, which no 0.29.x ships.
- CUDA OOM (58%): v0.29 costs ~1-2 GiB more init headroom (peak activation
  +1.13 GiB and an up-front cudagraph reserve, measured A/B on identical
  A40s with the same model and env), flipping marginal models from "fits"
  to "OOM at init".

Releasing this as v2.27.1 makes the Hub default safe again while the
fail-fast work lands separately; re-bump to a vLLM release that carries
vllm#56092 with huggingface_hub >= 1.30.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@lukepiette
lukepiette merged commit 32a1eed into runpod-workers:main Sep 21, 2026
6 of 9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants