Skip to content

[DO NOT MERGE] CI mirror of #351 (fork runs lack secrets) - #352

Closed
velaraptor-runpod wants to merge 2 commits into
mainfrom
pr-351
Closed

velaraptor-runpod wants to merge 2 commits into
mainfrom
pr-351

Conversation

@velaraptor-runpod

Copy link
Copy Markdown
Contributor

Internal branch copy of #351 (307752b) so secrets-gated jobs (Docker build, test-main) run with credentials. Close once checks pass; review and merge happen on #351.

ashwinsreedhar28 and others added 2 commits October 1, 2026 22:33
The image keeps the HF cache under BASE_PATH (/runpod-volume) but left
VLLM_CACHE_ROOT at ~/.cache/vllm, so every cold start recompiled the model
even with a network volume attached. Point it at $BASE_PATH/vllm-cache.
Without a volume BASE_PATH is a plain container directory, as before.

Docs: BASE_PATH is a build arg, so setting it on an endpoint moves neither
cache; say so in docs/configuration.md and the Hub field description.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JS2Me66WkkvXnWKiASyDpV
vLLM doesn't handle write errors on most compile-cache writes, so with the
cache on the volume a full or read-only volume would fail the start where
compiling in the container used to work:
- before launch, src/compile_cache.py checks the directory takes a written
  and synced block and has 1 GiB free, else uses vLLM's default root for
  that start;
- if vLLM still runs out of disk (the volume can fill during the weight
  download), main.py relaunches once with the cache in the container, like
  the existing revision and OOM relaunches, and before the OOM one.
The out-of-disk message now names the volume and vllm-cache, and quota
errors count as out of disk.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JS2Me66WkkvXnWKiASyDpV
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants