Repository navigation
Recover vanished HF revisions at startup and classify them correctly - #343
Conversation
|
@owencramer-runpod Dropped the 🤖 Generated with Claude Code |
Since v0.28 vLLM resolves every HF ref to an exact commit hash at launch (upstream's Artifact Pin Decay fix), so a force-pushed repo invalidates the pin and the boot dies with RevisionNotFoundError — 88% of failed v2.27.0 endpoints. The pre-flight catches a *configured* revision that never existed; this handles one that vanished after pre-flight, or a pin vLLM resolved itself. - main.py: relaunch vLLM once when startup died on a vanished revision — the one failure where retrying helps, since each launch re-resolves the pin; explicit *_REVISION env pins are dropped (loudly) for the retry. - startup_errors.py: classify RevisionNotFoundError before the generic 404 (it was misreported as "repository not found"); OOM advice now names the v0.29 recovery knobs (MAX_NUM_BATCHED_TOKENS=8192, KV_CACHE_DTYPE=auto on Ampere/Ada where fp8 KV forces FlashInfer's extra workspace). 79 tests pass (7 new). Nothing runs on the happy-path boot; the relaunch only executes after a startup failure that was previously fatal. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
0cb450b to
b6d63d6
Compare
|
Rebased onto main now that #342 and #344 are merged — rebuilt as a single commit on top of the pre-flight. The relaunch now composes with the pre-flight cleanly: the pre-flight catches a configured revision that never existed (fails in seconds, vLLM never launches); this PR handles a revision that vanished after pre-flight passed, or a pin vLLM resolved itself mid-boot. 79 tests pass (72 from main + 7 new). Ready for review. 🤖 Generated with Claude Code |
Promptless documentation updates
|
Problem
Hub template v2.27.0 (bundles vLLM v0.29.0) fails ~31% of first requests vs ~10–11% on v2.26.0 (vLLM 0.28.0). Worker-log analysis of 50 failed v2.27.0 endpoints shows 88% "Revision Not Found": since v0.28 vLLM resolves every HF ref to an exact commit hash at launch (upstream's fix for Artifact Pin Decay, GHSA-3ww4-5jv9-j5gm — so upstream won't restore main-fallback). A force-pushed repo invalidates the pinned hash →
RevisionNotFoundError. Worse, v0.29.0 ships huggingface_hub 1.28.0, which leaks a resolved pin across repos (model repo's hash applied to the tokenizer/weights repo → guaranteed 404); fixed upstream only in vllm #56092 + hf_hub ≥ 1.30. The engine rollback for that class is #344; this PR is the worker-side hardening that helps on every engine version.What this PR does
main.py— one relaunch for vanished revisions. When startup dies onRevisionNotFoundError, relaunchvllm serveonce: every launch re-resolves the pin, so this genuinely recovers the force-push case. ExplicitMODEL_REVISION/TOKENIZER_REVISION/CODE_REVISIONpins are dropped for the retry, loudly logged — serving the current branch beats crash-looping on a commit that no longer exists. A second failure answers jobs with a clear revision-specific message. Never a third attempt.startup_errors.py—RevisionNotFoundErrorclassified before the generic 404 (it previously matched "404 Client Error" and was misreported as repository-not-found, sending users to debug the wrong thing); OOM advice now names the v0.29-specific knobs (MAX_NUM_BATCHED_TOKENS=8192— the default doubled in v0.28;KV_CACHE_DTYPE=auto— fp8 KV forces FlashInfer on Ampere/Ada, which needs extra workspace, vllm #55509).Nothing runs on the happy-path boot: the relaunch logic only executes after a startup failure that would previously have been fatal.
Deliberately out of scope
HF_TOKENhub.json input: dropped per review — the console already collects a token contextually when a gated repo is selected.Tests
58 passed(7 new): revision classification & precedence over the generic 404, relaunch-once semantics (stale pins dropped, never a third attempt, OOM doesn't relaunch), v0.29 OOM knobs present.🤖 Generated with Claude Code