Repository navigation
fix: persist vLLM's torch.compile cache on the network volume - #351
Merged
velaraptor-runpod merged 3 commits intoOct 6, 2026
Merged
velaraptor-runpod merged 3 commits into
velaraptor-runpod merged 3 commits into
Conversation
The image keeps the HF cache under BASE_PATH (/runpod-volume) but left VLLM_CACHE_ROOT at ~/.cache/vllm, so every cold start recompiled the model even with a network volume attached. Point it at $BASE_PATH/vllm-cache. Without a volume BASE_PATH is a plain container directory, as before. Docs: BASE_PATH is a build arg, so setting it on an endpoint moves neither cache; say so in docs/configuration.md and the Hub field description. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JS2Me66WkkvXnWKiASyDpV
vLLM doesn't handle write errors on most compile-cache writes, so with the cache on the volume a full or read-only volume would fail the start where compiling in the container used to work: - before launch, src/compile_cache.py checks the directory takes a written and synced block and has 1 GiB free, else uses vLLM's default root for that start; - if vLLM still runs out of disk (the volume can fill during the weight download), main.py relaunches once with the cache in the container, like the existing revision and OOM relaunches, and before the OOM one. The out-of-disk message now names the volume and vllm-cache, and quota errors count as out of disk. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JS2Me66WkkvXnWKiASyDpV
velaraptor-runpod
self-requested a review
October 5, 2026 22:31
velaraptor-runpod
requested changes
Oct 5, 2026
Review feedback on runpod-workers#351: fall_back() returned False silently when VLLM_CACHE_ROOT and vLLM's default root are on the same filesystem. Log why nothing moved, and assert the warning in the existing test. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019HvA3d6MGzMtftEUMhRikc
velaraptor-runpod
approved these changes
Oct 6, 2026
Contributor
|
ty for the pr @ashwinsreedhar28 we will make the release on our hub on thursday! meantime main will be building the docker image and you can start using right away as soon as it is done building! |
Promptless documentation updates
|
Contributor
Author
thanks @velaraptor-runpod, appreciate the quick review! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The image puts the HF cache under
BASE_PATH(/runpod-volume) so model downloads survive cold starts.VLLM_CACHE_ROOT, though, stays at its default~/.cache/vllm, inside the container. So even with a network volume attached, every cold start compiles the model from an empty cache: 29 to 58 s of torch.compile on an RTX 4090 for Qwen3-8B in the runs below, and 32.6 s of a 154 s cold start in an earlier worker log.This sets
VLLM_CACHE_ROOT="${BASE_PATH}/vllm-cache"in the DockerfileENVblock, next toHF_HOME. The first start on a volume compiles and saves; later starts load the artifact.src/compile_cache.py, one check and one relaunch branch insrc/main.py, and anout_of_disk()predicate and a reworded message insrc/startup_errors.py. Both guards are needed: the pre-launch check avoids a failed launch on every cold start once the volume is full, and the relaunch covers the start that fills it.VLLM_CACHE_ROOTback at vLLM's default root (in the container) for that start. The synced write catches an exhausted quota even ifdfreports the storage pool's free space; 1 GiB is a heuristic.main.pyalready uses for a vanished revision and for OOM, and runs before the OOM relaunch, which would turn compilation off. It also fires when the weights themselves don't fit on the volume; that costs one quick extra launch.vllm-cache.BASE_PATHis a plain container directory (where the HF cache already goes), so the cache is lost on exit, same as today.--build-arg BASE_PATH=/models): the cache lands in/models/vllm-cacheinside the container; like today's/root/.cache/vllm, it is lost on exit.XDG_CACHE_HOME. Anyone who set it to move vLLM's cache is now overridden by the image'sVLLM_CACHE_ROOT.VLLM_CACHE_ROOT=/root/.cache/vllmon the endpoint.If you'd rather take only the one-line
ENVchange plus the docs and handle the full-volume case separately, I'm happy to cut the PR down to that.Checked in vLLM v0.30.0 / torch 2.13.0 (the pinned base image)
VLLM_CACHE_ROOTholds vLLM's torch.compile cache.VLLM_CACHE_ROOT/torch_compile_cache/torch_aot_compile/{hash}/rank_{r}_{dp}/model.TORCHINDUCTOR_CACHE_DIRis set to{hash}/inductor_cachebefore both load and compile (decorators.py#L527-L567).TRITON_CACHE_DIRunder it when it builds its first kernel (triton_heuristics.py#L492-L495).VLLM_USE_AOT_COMPILE=0): the standalone-compile artifacts go toVLLM_CACHE_ROOT/torch_compile_cache/{hash}/…(backends.py#L1058-L1079, compiler_interface.py#L388-L414). Inductor's and Triton's working caches stay in/tmpthere.TRITON_CACHE_DIR(~/.triton/cache), FlashInfer kernels that aren't in the image's prebuiltflashinfer-jit-cache, and CUDA graphs.get_inductor_factors()whenVLLM_USE_MEGA_AOT_ARTIFACTis on (caching.py#L573-L589). That's the default on torch ≥ 2.12 (envs.py#L379-L386), and the image ships torch 2.13.0.CacheBase.get_system()(compiler_interface.py#L170-L190): the GPU name plus the CUDA and Triton builds (codecache.py#L287-L311).VLLM_CACHE_ROOTis incompile_factors()' ignore list, envs.py#L2251-L2264), so moving the root invalidates nothing.VLLM_USE_MEGA_AOT_ARTIFACT=0(or an image built from a vLLM whose torch is 2.10–2.11), the AOT key has no GPU factor, so GPU types on one volume share one artifact path. An explicitcompilation_config.cache_diris used without any hash.Compiling model again due to a load failure, recompiles, and the new save overwrites it (decorators.py#L317-L330).VLLM_FORCE_AOT_LOAD=1turns that into a startup error.{path}.{pid}.tmpandos.replaces it into place (decorators.py#L711-L717)..tmpfile, and the next start recompiles.modelfile; the next start fails to load it, recompiles and overwrites it.$BASE_PATH/vllm-cacheis deleted. This happens withVLLM_USE_AOT_COMPILE=0(also forced byVLLM_BATCH_INVARIANT=1) when the cache index is truncated or an artifact it lists is missing (backends.py#L185-L221, #L250-L252). It also happens withVLLM_SKIP_P2P_CHECK=0when the P2P cache file is truncated.Other things vLLM 0.30 keeps under
VLLM_CACHE_ROOT:VLLM_SKIP_P2P_CHECK=0(default 1). With that set, a multi-GPU endpoint would trust another host's probe result.Evidence (Oct 5, 2026: RTX 4090 "24 GB PRO", Qwen/Qwen3-8B,
MAX_MODEL_LEN=4096, 30 GB network volume in EU-RO-1, FlashBoot off, CUDA 13.0 hosts, every sample a full cold boot)One endpoint built from this branch; the two arms alternated by adding and removing
VLLM_CACHE_ROOT=/root/.cache/vllmon the endpoint (the control, i.e. today's behavior). Both arms read the weights from the volume's HF cache. Times are from the worker log; "launch → healthy" is the wrapper'svllm servelaunch to itsvLLM is healthyline, so it excludes Runpod scheduling, the image pull and the SDK fitness checks. Runpod'sdelayTimeis shown for completeness but was dominated by waits for a 4090 on Low availability (86 to 860 s), not by the boot.delayTime/root/.cache/vllm)* s01 read the weights from the page cache right after downloading them; s02 to s06 read them cold off the FUSE volume (about 25 s), the same on both arms.
Directly load AOT compilation from path /runpod-volume/vllm-cache/torch_compile_cache/torch_aot_compile/…; the controls logUsing cache directory: /root/.cache/vllm/…and compile from scratch.collected artifacts: 37 entries, 3 artifacts, 6053782 bytes).xdzgazvaixvhk7) that wrote the cache. The control s03 ran on a different host (u0edpx1u3h235n). The cache lives on the network volume, not the host, so a cross-host hit should behave the same; I haven't measured one yet. The full-volume and cached-model fallbacks are covered by the unit tests below and not yet exercised on Serverless.Changes
Dockerfile:VLLM_CACHE_ROOT="${BASE_PATH}/vllm-cache", with a comment.src/compile_cache.py: the pre-launch check and the fallback. Pure Python, no vLLM import.src/main.py: calls the check before launch, and relaunches once on out-of-disk with the cache moved.src/startup_errors.py:out_of_disk()predicate.huggingface-cacheandvllm-cache.README.md,docs/configuration.md,docs/conventions.md:BASE_PATHnow holds both caches.BASE_PATHis a build arg, so setting it on an endpoint moves neither cache. That was already true for the HF cache..runpod/hub.json: theBASE_PATHfield's description now says it's build-time only and namesHF_HUB_CACHE/VLLM_CACHE_ROOTas the settings to change instead.tests/test_dockerfile_env.py: checks thatHF_HOMEandVLLM_CACHE_ROOTstay under${BASE_PATH}.tests/test_compile_cache.py:~expansion.main()runs the check before launching, relaunches once when the volume fills mid-start, never launches a third time on a second out-of-disk, puts the disk relaunch before the OOM relaunch, and skips the relaunch without a volume.tests/test_startup_errors.py: the new message, and quota errors matching.main()harnesses now unsetVLLM_CACHE_ROOT, so they don't probe a developer's own cache.python -m pytest tests -v: 142 passed. Each of the following has a test that fails when its code is removed: the relaunch and its ordering, the device check, the probe, the threshold, the opt-out, and~expansion.