Repository navigation
Auto-retry startup OOM once with a reduced memory footprint - #346
Merged
Merged
Conversation
A torch.OutOfMemoryError at boot was terminal even when the model would fit with a conservative setup — the dominant v2.27.0 failure class in warehouse data (~48% of failed first-request endpoints Sep 11-17, after the benign cudagraph-profiling INFO line is excluded from the count). vLLM v0.29 reserves cudagraph memory up front (kv pool = requested - non-KV - graph estimate; upstream vllm#57475, root fix vllm#51590 still open) and the batched-tokens default doubled in v0.28, so a boot that fit on v2.26.0 can OOM identically forever. Mirror the vanished-revision relaunch: when startup dies on any of the three memory wordings (torch OOM, no memory for cache blocks, KV cache too small), relaunch once with ENFORCE_EAGER=true (drops the up-front graph reserve) and MAX_NUM_BATCHED_TOKENS=8192 (halves peak activation). Each knob is only touched when it can still be relaxed: an explicit ENFORCE_EAGER stays overridden (loudly), a user-tightened token budget at or below 8192 is kept, and when nothing is left to relax the retry is skipped entirely. A second memory failure answers jobs with the cause, now stating the retry already ran and listing only the remaining real fixes (smaller/quantized checkpoint, larger GPU or TENSOR_PARALLEL_SIZE, lower MAX_MODEL_LEN, KV_CACHE_DTYPE=auto). The revision relaunch keeps its own single-use budget, so a boot may launch at most three times (revision retry then OOM retry); every other failure mode is answered or platform-retried exactly as before, and the happy path is untouched — the recovery only runs after a startup failure that would previously have been fatal. Tests: 92 passed (13 new/updated): one retry that can succeed, never a third launch, per-knob relaxation semantics, explicit-config respect, predicate coverage for all three memory wordings, retried-wording message content, revision+OOM budgets composing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
colin99d
approved these changes
Sep 25, 2026
Promptless documentation updates
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
CUDA OOM at init is the dominant worker-vllm first-request failure class — and it persisted after the v2.27.1 rollback . Drivers: v0.29 budgets the KV pool as
requested − non-KV − cudagraph_memory_estimate(graphs reserved up front — see vllm#57475; the proper fix, vllm#51590, is still open) and themax-num-batched-tokensdefault doubled in v0.28. A boot that OOMs on defaults is answered with advice… and then crash-loops on the exact same footprint. One relaxed retry can genuinely fit where the defaults could not (upstream's own issue shows--enforce-eagerbooting configs that OOM with graphs).What
A startup OOM now gets one reduced-footprint relaunch, mirroring the vanished-revision relaunch from #343:
startup_errors.memory_shortfall()— any of the three memory wordings (torch.OutOfMemoryError/ "CUDA out of memory", "No available memory for the cache blocks", "KV cache too small").ENFORCE_EAGER=true(drops the up-front graph reserve) andMAX_NUM_BATCHED_TOKENS=8192(halves peak activation).ENFORCE_EAGERvalue is overridden but loudly logged; a user-set token budget ≤ 8192 is kept; placeholderMAX_NUM_BATCHED_TOKENS=0counts as unset (matches args_builder's ZERO_MEANS_UNSET). When both knobs are already at/over their recovery value the retry is skipped as pointless and jobs get the existing advice immediately.TENSOR_PARALLEL_SIZE, lowerMAX_MODEL_LEN,KV_CACHE_DTYPE=autoon Ampere/Ada).Intent: on marginal models (24/48 GB cards) the endpoint comes up slowly-but-serving instead of crash-looping or dead. Degraded-mode intent is loudly logged (no response-metadata channel exists for it today).
Tests
pytest tests/— 92 passed (13 new/updated):tests/test_main_oom_retry.py(new): retry-that-succeeds, never-a-third, per-knob relaxation, explicit-config respect (ENFORCE_EAGER=true+ budget pre-tightened → no retry), KV wordings share the class, revision+OOM budgets compose (3 launches max).tests/test_startup_errors.py: predicate coverage for all three memory wordings; the post-retry message says what ran and omits the already-applied knobs.tests/test_main_revision_retry.py: its old "other fatal failures never relaunch" case used an OOM output — now a gated-repo error, matching the new policy.