Skip to content

fix: persist vLLM's torch.compile cache on the network volume - #351

Merged
velaraptor-runpod merged 3 commits into
runpod-workers:mainfrom
ashwinsreedhar28:fix/compile-cache-on-volume
Oct 6, 2026
Merged

velaraptor-runpod merged 3 commits into
runpod-workers:mainfrom
ashwinsreedhar28:fix/compile-cache-on-volume

Conversation

@ashwinsreedhar28

Copy link
Copy Markdown
Contributor

The image puts the HF cache under BASE_PATH (/runpod-volume) so model downloads survive cold starts. VLLM_CACHE_ROOT, though, stays at its default ~/.cache/vllm, inside the container. So even with a network volume attached, every cold start compiles the model from an empty cache: 29 to 58 s of torch.compile on an RTX 4090 for Qwen3-8B in the runs below, and 32.6 s of a 154 s cold start in an earlier worker log.

This sets VLLM_CACHE_ROOT="${BASE_PATH}/vllm-cache" in the Dockerfile ENV block, next to HF_HOME. The first start on a volume compiles and saves; later starts load the artifact.

  • Full or read-only volume. vLLM doesn't handle write errors on most compile-cache writes. Without a guard, a volume that is full (the cache is never pruned) or read-only would fail startup where compiling in the container used to work. This is the only reason the PR has runtime code rather than just the env var; it's kept to src/compile_cache.py, one check and one relaunch branch in src/main.py, and an out_of_disk() predicate and a reworded message in src/startup_errors.py. Both guards are needed: the pre-launch check avoids a failed launch on every cold start once the volume is full, and the relaunch covers the start that fills it.
    • Before launching vLLM, the worker checks that the directory takes a written and synced 64 KiB block and has at least 1 GiB free. If not, it logs why and points VLLM_CACHE_ROOT back at vLLM's default root (in the container) for that start. The synced write catches an exhausted quota even if df reports the storage pool's free space; 1 GiB is a heuristic.
    • During the start. The check runs before vLLM downloads the weights onto the same volume, so if vLLM still runs out of disk or quota, the worker relaunches once with the cache in the container. This follows the same one-relaunch pattern main.py already uses for a vanished revision and for OOM, and runs before the OOM relaunch, which would turn compilation off. It also fires when the weights themselves don't fit on the volume; that costs one quick extra launch.
    • Same disk. Moving the cache can't help when root and fallback are on the same disk, so the worker skips both the pre-launch fallback and the relaunch. The out-of-disk message now names the volume and vllm-cache.
  • No volume. BASE_PATH is a plain container directory (where the HF cache already goes), so the cache is lost on exit, same as today.
  • Baked-model images (--build-arg BASE_PATH=/models): the cache lands in /models/vllm-cache inside the container; like today's /root/.cache/vllm, it is lost on exit.
  • XDG_CACHE_HOME. Anyone who set it to move vLLM's cache is now overridden by the image's VLLM_CACHE_ROOT.
  • Opt out: set VLLM_CACHE_ROOT=/root/.cache/vllm on the endpoint.

If you'd rather take only the one-line ENV change plus the docs and handle the full-volume case separately, I'm happy to cut the PR down to that.

Checked in vLLM v0.30.0 / torch 2.13.0 (the pinned base image)

  1. VLLM_CACHE_ROOT holds vLLM's torch.compile cache.
    • AOT path, the default on torch ≥ 2.10 (envs.py#L367-L376).
      • The artifact goes to VLLM_CACHE_ROOT/torch_compile_cache/torch_aot_compile/{hash}/rank_{r}_{dp}/model.
      • TORCHINDUCTOR_CACHE_DIR is set to {hash}/inductor_cache before both load and compile (decorators.py#L527-L567).
      • Inductor then sets TRITON_CACHE_DIR under it when it builds its first kernel (triton_heuristics.py#L492-L495).
    • Non-AOT path (VLLM_USE_AOT_COMPILE=0): the standalone-compile artifacts go to VLLM_CACHE_ROOT/torch_compile_cache/{hash}/… (backends.py#L1058-L1079, compiler_interface.py#L388-L414). Inductor's and Triton's working caches stay in /tmp there.
    • Still rebuilt every cold start: vLLM's own Triton kernels compiled before Inductor sets TRITON_CACHE_DIR (~/.triton/cache), FlashInfer kernels that aren't in the image's prebuilt flashinfer-jit-cache, and CUDA graphs.
  2. With this image's defaults the cache key includes the GPU, so GPU types sharing a volume get separate entries.
    • The AOT hash adds get_inductor_factors() when VLLM_USE_MEGA_AOT_ARTIFACT is on (caching.py#L573-L589). That's the default on torch ≥ 2.12 (envs.py#L379-L386), and the image ships torch 2.13.0.
    • Those factors start with CacheBase.get_system() (compiler_interface.py#L170-L190): the GPU name plus the CUDA and Triton builds (codecache.py#L287-L311).
    • The path itself isn't a factor (VLLM_CACHE_ROOT is in compile_factors()' ignore list, envs.py#L2251-L2264), so moving the root invalidates nothing.
    • Two settings remove the GPU from the key. With VLLM_USE_MEGA_AOT_ARTIFACT=0 (or an image built from a vLLM whose torch is 2.10–2.11), the AOT key has no GPU factor, so GPU types on one volume share one artifact path. An explicit compilation_config.cache_dir is used without any hash.
  3. On the default path, a bad artifact costs a recompile.
    • The AOT load catches any exception. A corrupt or incompatible file logs Compiling model again due to a load failure, recompiles, and the new save overwrites it (decorators.py#L317-L330). VLLM_FORCE_AOT_LOAD=1 turns that into a startup error.
    • Inductor skips unreadable FX-graph cache entries with a warning (codecache.py#L1711-L1719).
    • The save writes {path}.{pid}.tmp and os.replaces it into place (decorators.py#L711-L717).
      • A worker killed mid-save leaves a stray .tmp file, and the next start recompiles.
      • PIDs are per container, so two workers saving the same config at once can pick the same temp name and interleave writes. That can leave a corrupt model file; the next start fails to load it, recompiles and overwrites it.
    • Off the default path, a bad file fails every cold start until $BASE_PATH/vllm-cache is deleted. This happens with VLLM_USE_AOT_COMPILE=0 (also forced by VLLM_BATCH_INVARIANT=1) when the cache index is truncated or an artifact it lists is missing (backends.py#L185-L221, #L250-L252). It also happens with VLLM_SKIP_P2P_CHECK=0 when the P2P cache file is truncated.

Other things vLLM 0.30 keeps under VLLM_CACHE_ROOT:

  • Safe to share: the model-info cache (keyed by the model module's source), the opt-in startup plan (keyed by GPU name, memory and capability), and the DeepGEMM JIT cache.
  • FlashInfer autotune (SM90+): keyed by compute capability, not GPU name, so H100 variants share tactics. That affects performance only.
  • GPU P2P probe cache: keyed by GPU indices only, and read only with VLLM_SKIP_P2P_CHECK=0 (default 1). With that set, a multi-GPU endpoint would trust another host's probe result.

Evidence (Oct 5, 2026: RTX 4090 "24 GB PRO", Qwen/Qwen3-8B, MAX_MODEL_LEN=4096, 30 GB network volume in EU-RO-1, FlashBoot off, CUDA 13.0 hosts, every sample a full cold boot)

One endpoint built from this branch; the two arms alternated by adding and removing VLLM_CACHE_ROOT=/root/.cache/vllm on the endpoint (the control, i.e. today's behavior). Both arms read the weights from the volume's HF cache. Times are from the worker log; "launch → healthy" is the wrapper's vllm serve launch to its vLLM is healthy line, so it excludes Runpod scheduling, the image pull and the SDK fitness checks. Runpod's delayTime is shown for completeness but was dominated by waits for a 4090 on Low availability (86 to 860 s), not by the boot.

sample arm torch.compile init engine weights off volume launch → healthy delayTime
s01 first start (downloads, compiles, saves) 58.0 s 87.6 s 2.8 s* 142 s 860 s
s02 this PR (cache on volume) 1.4 s 19.8 s 27.0 s 71 s 159 s
s03 control (/root/.cache/vllm) 40.7 s 63.5 s 28.4 s 133 s 227 s
s04 this PR 1.1 s 20.6 s 24.7 s 70 s 141 s
s05 control 29.2 s 48.1 s 26.1 s 108 s 771 s
s06 this PR 1.0 s 19.2 s 24.7 s 68 s 86 s

* s01 read the weights from the page cache right after downloading them; s02 to s06 read them cold off the FUSE volume (about 25 s), the same on both arms.

  • The cache hits log Directly load AOT compilation from path /runpod-volume/vllm-cache/torch_compile_cache/torch_aot_compile/…; the controls log Using cache directory: /root/.cache/vllm/… and compile from scratch.
  • Compile from scratch varied 29 to 58 s across the three cold compiles (host CPU load); the three cache hits were 1.0 to 1.4 s. Launch → healthy: 68 to 71 s with the PR vs 108 to 133 s for the controls, a saving of about 40 to 60 s per cold start.
  • The compile artifacts on the volume are about 6 MB (collected artifacts: 37 entries, 3 artifacts, 6053782 bytes).
  • Caveat: all three cache hits landed on the same host (xdzgazvaixvhk7) that wrote the cache. The control s03 ran on a different host (u0edpx1u3h235n). The cache lives on the network volume, not the host, so a cross-host hit should behave the same; I haven't measured one yet. The full-volume and cached-model fallbacks are covered by the unit tests below and not yet exercised on Serverless.

Changes

  • Dockerfile: VLLM_CACHE_ROOT="${BASE_PATH}/vllm-cache", with a comment.
  • src/compile_cache.py: the pre-launch check and the fallback. Pure Python, no vLLM import.
  • src/main.py: calls the check before launch, and relaunches once on out-of-disk with the cache moved.
  • src/startup_errors.py:
    • New out_of_disk() predicate.
    • Disk-quota errors now match, not only ENOSPC.
    • The out-of-disk message names the volume, huggingface-cache and vllm-cache.
  • README.md, docs/configuration.md, docs/conventions.md:
    • BASE_PATH now holds both caches.
    • The fallback, the opt-out, and that the cache is never pruned. It's safe to delete while no worker is starting.
    • BASE_PATH is a build arg, so setting it on an endpoint moves neither cache. That was already true for the HF cache.
  • .runpod/hub.json: the BASE_PATH field's description now says it's build-time only and names HF_HUB_CACHE / VLLM_CACHE_ROOT as the settings to change instead.
  • tests/test_dockerfile_env.py: checks that HF_HOME and VLLM_CACHE_ROOT stay under ${BASE_PATH}.
  • tests/test_compile_cache.py:
    • Fallback cases: an uncreatable root, an existing read-only root, an exhausted quota, a nearly full volume.
    • A failing device lookup still falls back.
    • Cases where it must not move the cache: no volume, the opt-out, and ~ expansion.
    • main() runs the check before launching, relaunches once when the volume fills mid-start, never launches a third time on a second out-of-disk, puts the disk relaunch before the OOM relaunch, and skips the relaunch without a volume.
  • tests/test_startup_errors.py: the new message, and quota errors matching.
  • The existing main() harnesses now unset VLLM_CACHE_ROOT, so they don't probe a developer's own cache.

python -m pytest tests -v: 142 passed. Each of the following has a test that fails when its code is removed: the relaunch and its ordering, the device check, the probe, the threshold, the opt-out, and ~ expansion.

ashwinsreedhar28 and others added 2 commits October 1, 2026 22:33
The image keeps the HF cache under BASE_PATH (/runpod-volume) but left
VLLM_CACHE_ROOT at ~/.cache/vllm, so every cold start recompiled the model
even with a network volume attached. Point it at $BASE_PATH/vllm-cache.
Without a volume BASE_PATH is a plain container directory, as before.

Docs: BASE_PATH is a build arg, so setting it on an endpoint moves neither
cache; say so in docs/configuration.md and the Hub field description.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JS2Me66WkkvXnWKiASyDpV
vLLM doesn't handle write errors on most compile-cache writes, so with the
cache on the volume a full or read-only volume would fail the start where
compiling in the container used to work:
- before launch, src/compile_cache.py checks the directory takes a written
  and synced block and has 1 GiB free, else uses vLLM's default root for
  that start;
- if vLLM still runs out of disk (the volume can fill during the weight
  download), main.py relaunches once with the cache in the container, like
  the existing revision and OOM relaunches, and before the OOM one.
The out-of-disk message now names the volume and vllm-cache, and quota
errors count as out of disk.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JS2Me66WkkvXnWKiASyDpV
@velaraptor-runpod
velaraptor-runpod self-requested a review October 5, 2026 22:31
Comment thread src/compile_cache.py
Review feedback on runpod-workers#351: fall_back() returned False silently when
VLLM_CACHE_ROOT and vLLM's default root are on the same filesystem.
Log why nothing moved, and assert the warning in the existing test.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HvA3d6MGzMtftEUMhRikc
@velaraptor-runpod
velaraptor-runpod merged commit 2a983f1 into runpod-workers:main Oct 6, 2026
1 check passed
@velaraptor-runpod

Copy link
Copy Markdown
Contributor

ty for the pr @ashwinsreedhar28

we will make the release on our hub on thursday! meantime main will be building the docker image and you can start using right away as soon as it is done building!

@promptless

promptless Bot commented Oct 6, 2026

Copy link
Copy Markdown

Promptless documentation updates

@ashwinsreedhar28

Copy link
Copy Markdown
Contributor Author

ty for the pr @ashwinsreedhar28

we will make the release on our hub on thursday! meantime main will be building the docker image and you can start using right away as soon as it is done building!

thanks @velaraptor-runpod, appreciate the quick review!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants