Repository navigation
fix: raise default VLLM_STARTUP_TIMEOUT to 1800s and expose it in the hub env schema - #354
Merged
Merged
Conversation
JessicaGarson
approved these changes
Oct 7, 2026
Promptless documentation updates
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
src/main.pyfrom 1200s to 1800s (VLLM_STARTUP_TIMEOUToverride unchanged).VLLM_STARTUP_TIMEOUTto the.runpod/hub.jsonenv schema (advanced, default 1800) so hub/console deploys surface it; today the knob exists in code but is invisible to users.docs/configuration.mdin sync.Why
On 2026-10-06 a 4×H200 worker serving DeepSeek-V4.1-Flash (475 GB, fp8_ds_mla) had its weight load stretched to ~13 min by same-host contention from overspawned sibling workers. The 1200s watchdog fired and SIGTERM'd the vLLM process 13 seconds before it printed "Application startup complete", forcing a full container restart and another multi-minute engine bring-up. For 400 GB+ models, 1200s leaves no headroom for variance in page-cache warmth, host contention, or disk pressure — and killing an actively-loading backend wastes the entire load it just paid for.
1800s covers the observed worst case (789s weight load + graph capture + warmup) with margin, while still bounding genuinely stuck starts.
Not in this PR
Follow-up worth a separate issue: make the watchdog progress-aware (don't SIGTERM a backend still actively loading weights) instead of purely wall-clock.
Testing
pytest tests— 142 passed.hub.jsonparses and matches the sibling entry shape.