Skip to content

Pre-flight the model reference before booting vLLM - #342

Merged
lukepiette merged 1 commit into
runpod-workers:mainfrom
lukepiette:preflight-model-validation
Sep 21, 2026
Merged

lukepiette merged 1 commit into
runpod-workers:mainfrom
lukepiette:preflight-model-validation

Conversation

@lukepiette

Copy link
Copy Markdown
Contributor

Problem

When vllm serve downloads the weights itself (no baked-in model), a typo'd MODEL_NAME, a gated model without HF_TOKEN, or a bad MODEL_REVISION only surfaces after the full cold start (~19 min), when the process dies and startup_errors.classify() parses the crash output. The user waits ~20 minutes to learn their config was wrong. This class of error is ~20% of worker-vllm's first-request failures.

Before / After

Before After
Typo'd model id ~19 min download attempt → crash → classified message ~0.1 s: … was not found on Hugging Face. Check MODEL_NAME for typos (it must be the full \org/repo` id) …`
Gated model, no token ~19 min → crash → classified message ~0.2 s: … is gated or private on Hugging Face … Set HF_TOKEN to a token whose account has accepted the model's license …
Bad MODEL_REVISION ~19 min → crash, generic ~0.1 s: MODEL_REVISION=… does not exist for … Check the model page for the branch, tag, or commit hash …
Valid public model boots boots unchanged (one metadata call added, ~0.1 s)
Transient network error retried still retried — never misreported as not-found; boot proceeds and vLLM makes its own attempt

How

  • New src/model_preflight.py: a metadata-only HfApi().model_info() call (no weights downloaded) validating the exact values args_builder.py hands to vLLM (MODEL_NAME, MODEL_REVISION, HF_TOKEN). Gated repos serve their metadata publicly, so a truthy gated flag is followed by auth_check() to learn whether this token may actually download the files.
  • main() runs the pre-flight before start_vllm(). On a definitive failure it never launches vLLM and follows the existing no-crash-loop pattern: startup_error is set and every job is answered with the cause.
  • Skipped entirely for baked-in models (Option 2 sets HF_HUB_OFFLINE/TRANSFORMERS_OFFLINE), local paths, non-HF sources, and MODEL/VLLM_CONFIG_FILE deploys.
  • The gated / not-found wording is factored into shared helpers in startup_errors.py, so the fast path and the post-crash path emit identical messages (asserted by a test).

Validation

  • pytest tests/ — 72 passed (67 existing + new tests/test_model_preflight.py covering nonexistent repo, gated±token, bad revision, 401/403, transient 5xx/429/timeout/DNS → proceed, and every skip case).
  • Live dry-run against the real Hub (no token, implicit token disabled):
    --- nonexistent repo (0.12s)  → "definitely-not-a-real-org/nope-123 was not found on Hugging Face. …"
    --- valid public model (0.09s) → OK, boot proceeds
    --- gated, no token (0.17s)   → "meta-llama/Llama-3.1-8B-Instruct is gated or private … Set HF_TOKEN …"
    --- bad revision (0.08s)      → "MODEL_REVISION=no-such-rev does not exist for Qwen/Qwen2.5-0.5B-Instruct …"
    

🤖 Generated with Claude Code

When vllm serve downloads the weights itself, a typo'd MODEL_NAME, a
gated model without HF_TOKEN, or a bad MODEL_REVISION only surfaced
after the full ~20-minute cold start, when startup_errors.classify
parsed the crash output. One metadata-only call to the HF Hub answers
the same question in seconds, so main() now asks it before launching
vLLM and answers jobs with the cause (same no-crash-loop rule as the
post-crash path, and the same wording, now shared via helpers in
startup_errors.py).

Only definitive Hub answers fail the boot: offline/baked-in models,
local paths, non-HF sources, and any transient network error all let
the boot proceed, so a blip is never misreported as "not found".
Gated repos serve their metadata publicly, so a truthy `gated` flag on
model_info is followed by auth_check to learn whether this token may
actually download the weights.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@lukepiette

Copy link
Copy Markdown
Contributor Author

Real-GPU validation (Runpod secure-cloud A40, CUDA 13.0 host, vllm 0.29.0 — the version this repo's Dockerfile pins):

  • Unit suite on the GPU host: 72 passed.
  • Pre-flight against the live HF API from the datacenter: nonexistent repo → clear not-found error in 0.10s; gated repo without token → HF_TOKEN error in 0.06s; bad MODEL_REVISION → revision error in 0.03s; valid public model → passes in 0.03s.
  • Bad model through the real entrypoint (python src/main.py): Model pre-flight failed; answering jobs with the cause instead of starting vLLM: definitely-not-a-real-org/nope-123 was not found… logged immediately, and vllm serve was never launched — versus the previous behavior of downloading/booting for ~19 min before dying.
  • Valid model (Qwen/Qwen2.5-0.5B-Instruct) through the same entrypoint: pre-flight passed, vllm serve launched, engine initialized on the A40, /health 200 and vLLM is healthy after 138s, /v1/models serving the model — the happy path is unchanged.

🤖 Generated with Claude Code

@lukepiette
lukepiette requested a review from OWCramer September 21, 2026 21:33
@lukepiette
lukepiette marked this pull request as ready for review September 21, 2026 21:39
@lukepiette
lukepiette merged commit 4262e66 into runpod-workers:main Sep 21, 2026
6 of 9 checks passed
@promptless

promptless Bot commented Sep 21, 2026

Copy link
Copy Markdown

Promptless documentation updates

  • Document vLLM worker startup error handling in Serverless troubleshooting updates the Serverless troubleshooting page's vLLM section for this change: a vanished/nonexistent Hugging Face revision is now its own troubleshooting case (no longer reported as repository-not-found), documenting the worker's automatic one-time relaunch against the current default branch (dropping the MODEL_REVISION/TOKENIZER_REVISION/CODE_REVISION pins), and adds the vLLM 0.29 OOM knobs MAX_NUM_BATCHED_TOKENS=8192 and KV_CACHE_DTYPE=auto.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants