Skip to content

Auto-retry startup OOM once with a reduced memory footprint - #346

Merged
velaraptor-runpod merged 1 commit into
mainfrom
feature/oom-startup-retry
Sep 25, 2026
Merged

velaraptor-runpod merged 1 commit into
mainfrom
feature/oom-startup-retry

Conversation

@velaraptor-runpod

@velaraptor-runpod velaraptor-runpod commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Why

CUDA OOM at init is the dominant worker-vllm first-request failure class — and it persisted after the v2.27.1 rollback . Drivers: v0.29 budgets the KV pool as requested − non-KV − cudagraph_memory_estimate (graphs reserved up front — see vllm#57475; the proper fix, vllm#51590, is still open) and the max-num-batched-tokens default doubled in v0.28. A boot that OOMs on defaults is answered with advice… and then crash-loops on the exact same footprint. One relaxed retry can genuinely fit where the defaults could not (upstream's own issue shows --enforce-eager booting configs that OOM with graphs).

What

A startup OOM now gets one reduced-footprint relaunch, mirroring the vanished-revision relaunch from #343:

  • Trigger: startup_errors.memory_shortfall() — any of the three memory wordings (torch.OutOfMemoryError / "CUDA out of memory", "No available memory for the cache blocks", "KV cache too small").
  • Relaunch env: ENFORCE_EAGER=true (drops the up-front graph reserve) and MAX_NUM_BATCHED_TOKENS=8192 (halves peak activation).
  • Per-knob relaxation, never blind override: an explicit ENFORCE_EAGER value is overridden but loudly logged; a user-set token budget ≤ 8192 is kept; placeholder MAX_NUM_BATCHED_TOKENS=0 counts as unset (matches args_builder's ZERO_MEANS_UNSET). When both knobs are already at/over their recovery value the retry is skipped as pointless and jobs get the existing advice immediately.
  • If the second launch OOMs too, jobs are answered with the cause — the message now states the retry already ran and lists only the remaining real fixes (smaller/quantized checkpoint, larger GPU or TENSOR_PARALLEL_SIZE, lower MAX_MODEL_LEN, KV_CACHE_DTYPE=auto on Ampere/Ada).
  • Bounded: at most 3 launches total (revision retry and OOM retry are independent single-use budgets). Every other failure mode and the happy path are byte-for-byte unchanged — the recovery code only runs after a startup failure that would previously have been terminal.

Intent: on marginal models (24/48 GB cards) the endpoint comes up slowly-but-serving instead of crash-looping or dead. Degraded-mode intent is loudly logged (no response-metadata channel exists for it today).

Tests

pytest tests/ — 92 passed (13 new/updated):

  • tests/test_main_oom_retry.py (new): retry-that-succeeds, never-a-third, per-knob relaxation, explicit-config respect (ENFORCE_EAGER=true + budget pre-tightened → no retry), KV wordings share the class, revision+OOM budgets compose (3 launches max).
  • tests/test_startup_errors.py: predicate coverage for all three memory wordings; the post-retry message says what ran and omits the already-applied knobs.
  • tests/test_main_revision_retry.py: its old "other fatal failures never relaunch" case used an OOM output — now a gated-repo error, matching the new policy.

A torch.OutOfMemoryError at boot was terminal even when the model would
fit with a conservative setup — the dominant v2.27.0 failure class in
warehouse data (~48% of failed first-request endpoints Sep 11-17, after
the benign cudagraph-profiling INFO line is excluded from the count).
vLLM v0.29 reserves cudagraph memory up front (kv pool = requested -
non-KV - graph estimate; upstream vllm#57475, root fix vllm#51590 still
open) and the batched-tokens default doubled in v0.28, so a boot that
fit on v2.26.0 can OOM identically forever.

Mirror the vanished-revision relaunch: when startup dies on any of the
three memory wordings (torch OOM, no memory for cache blocks, KV cache
too small), relaunch once with ENFORCE_EAGER=true (drops the up-front
graph reserve) and MAX_NUM_BATCHED_TOKENS=8192 (halves peak activation).
Each knob is only touched when it can still be relaxed: an explicit
ENFORCE_EAGER stays overridden (loudly), a user-tightened token budget
at or below 8192 is kept, and when nothing is left to relax the retry
is skipped entirely. A second memory failure answers jobs with the
cause, now stating the retry already ran and listing only the remaining
real fixes (smaller/quantized checkpoint, larger GPU or
TENSOR_PARALLEL_SIZE, lower MAX_MODEL_LEN, KV_CACHE_DTYPE=auto).

The revision relaunch keeps its own single-use budget, so a boot may
launch at most three times (revision retry then OOM retry); every other
failure mode is answered or platform-retried exactly as before, and the
happy path is untouched — the recovery only runs after a startup failure
that would previously have been fatal.

Tests: 92 passed (13 new/updated): one retry that can succeed, never a
third launch, per-knob relaxation semantics, explicit-config respect,
predicate coverage for all three memory wordings, retried-wording
message content, revision+OOM budgets composing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@velaraptor-runpod
velaraptor-runpod merged commit 16d6377 into main Sep 25, 2026
12 checks passed
@velaraptor-runpod
velaraptor-runpod deleted the feature/oom-startup-retry branch September 25, 2026 20:56
@promptless

promptless Bot commented Sep 25, 2026

Copy link
Copy Markdown

Promptless documentation updates

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants