Repository navigation
Pre-flight the model reference before booting vLLM - #342
Merged
lukepiette merged 1 commit intoSep 21, 2026
Merged
Conversation
When vllm serve downloads the weights itself, a typo'd MODEL_NAME, a gated model without HF_TOKEN, or a bad MODEL_REVISION only surfaced after the full ~20-minute cold start, when startup_errors.classify parsed the crash output. One metadata-only call to the HF Hub answers the same question in seconds, so main() now asks it before launching vLLM and answers jobs with the cause (same no-crash-loop rule as the post-crash path, and the same wording, now shared via helpers in startup_errors.py). Only definitive Hub answers fail the boot: offline/baked-in models, local paths, non-HF sources, and any transient network error all let the boot proceed, so a blip is never misreported as "not found". Gated repos serve their metadata publicly, so a truthy `gated` flag on model_info is followed by auth_check to learn whether this token may actually download the weights. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Contributor
Author
|
Real-GPU validation (Runpod secure-cloud A40, CUDA 13.0 host, vllm 0.29.0 — the version this repo's Dockerfile pins):
🤖 Generated with Claude Code |
lukepiette
marked this pull request as ready for review
September 21, 2026 21:39
OWCramer
approved these changes
Sep 21, 2026
Promptless documentation updates
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
When
vllm servedownloads the weights itself (no baked-in model), a typo'dMODEL_NAME, a gated model withoutHF_TOKEN, or a badMODEL_REVISIONonly surfaces after the full cold start (~19 min), when the process dies andstartup_errors.classify()parses the crash output. The user waits ~20 minutes to learn their config was wrong. This class of error is ~20% of worker-vllm's first-request failures.Before / After
… was not found on Hugging Face. Check MODEL_NAME for typos (it must be the full \org/repo` id) …`… is gated or private on Hugging Face … Set HF_TOKEN to a token whose account has accepted the model's license …MODEL_REVISIONMODEL_REVISION=… does not exist for … Check the model page for the branch, tag, or commit hash …How
src/model_preflight.py: a metadata-onlyHfApi().model_info()call (no weights downloaded) validating the exact valuesargs_builder.pyhands to vLLM (MODEL_NAME,MODEL_REVISION,HF_TOKEN). Gated repos serve their metadata publicly, so a truthygatedflag is followed byauth_check()to learn whether this token may actually download the files.main()runs the pre-flight beforestart_vllm(). On a definitive failure it never launches vLLM and follows the existing no-crash-loop pattern:startup_erroris set and every job is answered with the cause.HF_HUB_OFFLINE/TRANSFORMERS_OFFLINE), local paths, non-HF sources, andMODEL/VLLM_CONFIG_FILEdeploys.startup_errors.py, so the fast path and the post-crash path emit identical messages (asserted by a test).Validation
pytest tests/— 72 passed (67 existing + newtests/test_model_preflight.pycovering nonexistent repo, gated±token, bad revision, 401/403, transient 5xx/429/timeout/DNS → proceed, and every skip case).🤖 Generated with Claude Code