Repository navigation
Roll the engine back to vLLM v0.28.0 until the v0.29 first-request regressions are fixed - #344
Merged
Conversation
…gressions are fixed Template v2.27.0 (vLLM v0.29.0) fails ~31% of first requests against ~10-11% on v2.26.0 (v0.28.0). Two engine-side regressions dominate the 50 failed endpoints analysed: - "Revision Not Found" (88%): v0.29.0 bundles huggingface_hub 1.28.0, whose resolved-revision pin leaks across repos (model hash applied to the tokenizer/weights repo -> guaranteed 404); fixed upstream only in vllm#56092 + hf_hub >= 1.30, which no 0.29.x ships. - CUDA OOM (58%): v0.29 costs ~1-2 GiB more init headroom (peak activation +1.13 GiB and an up-front cudagraph reserve, measured A/B on identical A40s with the same model and env), flipping marginal models from "fits" to "OOM at init". Releasing this as v2.27.1 makes the Hub default safe again while the fail-fast work lands separately; re-bump to a vLLM release that carries vllm#56092 with huggingface_hub >= 1.30. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
OWCramer
approved these changes
Sep 21, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Hub template v2.27.0 (vLLM v0.29.0) fails ~31% of first requests vs ~10–11% on v2.26.0 (v0.28.0), measured Sep 11–17. The only functional change in v2.27.0 was the engine bump, and both dominant failure modes are engine-side:
RevisionNotFoundError. Properly fixed only in vllm#56092 + hf_hub ≥ 1.30, which no 0.29.x carries. (Hash pinning itself is deliberate upstream — the Artifact Pin Decay advisory — so only the leak, not the pinning, is fixable.)Qwen2.5-Coder-32B-AWQ), same env: v0.29 uses +1.13 GiB peak activation (1.48 → 2.61 GiB) and reserves estimated cudagraph memory up front (its own log: requested 0.95 ≡ effective 0.9247), shrinking available KV cache 5% and flipping marginal models from "fits" to "OOM at init".What
One-line engine pin
v0.29.0 → v0.28.0plus the README version strings. The v2.27.0-era wrapper improvements (startup-error classification, flag-parser compat) are kept —check_vllm_flags/args_buildersupport both parsers.Shipping
Merging this and cutting v2.27.1 makes the Hub default template safe again (validated live: the same model/env that boots on v2.26.0 boots on v0.28.0-based images; see the A/B in #343). The fail-fast hardening (revision relaunch, pre-download fit check, HF_TOKEN input) lands separately in #343 and is version-independent. Re-bump the engine once a vLLM release carries vllm#56092 with hf_hub ≥ 1.30.
Re-measure: after v2.27.1 is the Hub default, compare first-request success for endpoints created on v2.27.1 vs the v2.27.0 cohort over one week; target is back to ~10–11% failure.
🤖 Generated with Claude Code