Skip to content

Recover vanished HF revisions at startup and classify them correctly - #343

Merged
lukepiette merged 1 commit into
runpod-workers:mainfrom
lukepiette:v029-first-request-fixes
Sep 21, 2026
Merged

lukepiette merged 1 commit into
runpod-workers:mainfrom
lukepiette:v029-first-request-fixes

Conversation

@lukepiette

@lukepiette lukepiette commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Problem

Hub template v2.27.0 (bundles vLLM v0.29.0) fails ~31% of first requests vs ~10–11% on v2.26.0 (vLLM 0.28.0). Worker-log analysis of 50 failed v2.27.0 endpoints shows 88% "Revision Not Found": since v0.28 vLLM resolves every HF ref to an exact commit hash at launch (upstream's fix for Artifact Pin Decay, GHSA-3ww4-5jv9-j5gm — so upstream won't restore main-fallback). A force-pushed repo invalidates the pinned hash → RevisionNotFoundError. Worse, v0.29.0 ships huggingface_hub 1.28.0, which leaks a resolved pin across repos (model repo's hash applied to the tokenizer/weights repo → guaranteed 404); fixed upstream only in vllm #56092 + hf_hub ≥ 1.30. The engine rollback for that class is #344; this PR is the worker-side hardening that helps on every engine version.

What this PR does

  1. main.py — one relaunch for vanished revisions. When startup dies on RevisionNotFoundError, relaunch vllm serve once: every launch re-resolves the pin, so this genuinely recovers the force-push case. Explicit MODEL_REVISION/TOKENIZER_REVISION/CODE_REVISION pins are dropped for the retry, loudly logged — serving the current branch beats crash-looping on a commit that no longer exists. A second failure answers jobs with a clear revision-specific message. Never a third attempt.
  2. startup_errors.py — RevisionNotFoundError classified before the generic 404 (it previously matched "404 Client Error" and was misreported as repository-not-found, sending users to debug the wrong thing); OOM advice now names the v0.29-specific knobs (MAX_NUM_BATCHED_TOKENS=8192 — the default doubled in v0.28; KV_CACHE_DTYPE=auto — fp8 KV forces FlashInfer on Ampere/Ada, which needs extra workspace, vllm #55509).

Nothing runs on the happy-path boot: the relaunch logic only executes after a startup failure that would previously have been fatal.

Deliberately out of scope

Tests

58 passed (7 new): revision classification & precedence over the generic 404, relaunch-once semantics (stale pins dropped, never a third attempt, OOM doesn't relaunch), v0.29 OOM knobs present.

🤖 Generated with Claude Code

@lukepiette
lukepiette requested a review from OWCramer September 21, 2026 21:33
@lukepiette
lukepiette marked this pull request as ready for review September 21, 2026 21:39
@lukepiette lukepiette changed the title Fail fast on the v0.29 first-request killers: vanished revisions and oversized models Recover vanished HF revisions at startup and classify them correctly Sep 21, 2026
@lukepiette

Copy link
Copy Markdown
Contributor Author

@owencramer-runpod Dropped the hub.json HF_TOKEN input — you're right that main-ui already collects it contextually for gated repos, so this would have been an always-visible duplicate. The PR is now only the revision handling: the relaunch-once in main.py and the RevisionNotFoundError classification (plus the v0.29 OOM knobs in the advice text). Nothing on the happy-path boot. 58 tests pass. Ready for another look.

🤖 Generated with Claude Code

Since v0.28 vLLM resolves every HF ref to an exact commit hash at launch
(upstream's Artifact Pin Decay fix), so a force-pushed repo invalidates the
pin and the boot dies with RevisionNotFoundError — 88% of failed v2.27.0
endpoints. The pre-flight catches a *configured* revision that never
existed; this handles one that vanished after pre-flight, or a pin vLLM
resolved itself.

- main.py: relaunch vLLM once when startup died on a vanished revision —
  the one failure where retrying helps, since each launch re-resolves the
  pin; explicit *_REVISION env pins are dropped (loudly) for the retry.
- startup_errors.py: classify RevisionNotFoundError before the generic 404
  (it was misreported as "repository not found"); OOM advice now names the
  v0.29 recovery knobs (MAX_NUM_BATCHED_TOKENS=8192, KV_CACHE_DTYPE=auto on
  Ampere/Ada where fp8 KV forces FlashInfer's extra workspace).

79 tests pass (7 new). Nothing runs on the happy-path boot; the relaunch
only executes after a startup failure that was previously fatal.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@lukepiette
lukepiette force-pushed the v029-first-request-fixes branch from 0cb450b to b6d63d6 Compare September 21, 2026 22:33
@lukepiette

Copy link
Copy Markdown
Contributor Author

Rebased onto main now that #342 and #344 are merged — rebuilt as a single commit on top of the pre-flight. The relaunch now composes with the pre-flight cleanly: the pre-flight catches a configured revision that never existed (fails in seconds, vLLM never launches); this PR handles a revision that vanished after pre-flight passed, or a pin vLLM resolved itself mid-boot. 79 tests pass (72 from main + 7 new). Ready for review.

🤖 Generated with Claude Code

@lukepiette
lukepiette merged commit 0730f8d into runpod-workers:main Sep 21, 2026
6 of 8 checks passed
@promptless

promptless Bot commented Sep 21, 2026

Copy link
Copy Markdown

Promptless documentation updates

  • Document vLLM worker startup error handling in Serverless troubleshooting updates the Serverless troubleshooting page's vLLM section for this change: a vanished/nonexistent Hugging Face revision is now its own troubleshooting case (no longer reported as repository-not-found), documenting the worker's automatic one-time relaunch against the current default branch (dropping the MODEL_REVISION/TOKENIZER_REVISION/CODE_REVISION pins), and adds the vLLM 0.29 OOM knobs MAX_NUM_BATCHED_TOKENS=8192 and KV_CACHE_DTYPE=auto.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants