Skip to content
Closed
Show file tree
Hide file tree
Changes from 7 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
30 changes: 30 additions & 0 deletions benchmarks/patches/vllm/minimax_m3_nvfp4_marlin.patch
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
--- a/vllm/model_executor/layers/fused_moe/oracle/nvfp4.py
+++ b/vllm/model_executor/layers/fused_moe/oracle/nvfp4.py
@@ -176,2 +176,3 @@
NvFp4MoeBackend.FLASHINFER_TRTLLM,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe upstream if sm120 then marlin else flashinfer?

+ NvFp4MoeBackend.MARLIN,
}
--- a/vllm/model_executor/layers/fused_moe/experts/marlin_moe.py
+++ b/vllm/model_executor/layers/fused_moe/experts/marlin_moe.py
@@ -583,9 +583,12 @@
- # Gated-activation params (used by SWIGLUOAI_UNINTERLEAVE on packed w13).
- # silu == swigluoai with alpha=1, beta=0; configs that don't set these
- # (plain silu) fall back to the silu identity.
- self.gemm1_alpha = (
- quant_config.gemm1_alpha if quant_config.gemm1_alpha is not None else 1.0
- )
- self.gemm1_beta = (
- quant_config.gemm1_beta if quant_config.gemm1_beta is not None else 0.0
- )
+ if self.gemm1_clamp_limit is None:
+ self.gemm1_clamp_limit = moe_config.swiglu_limit
+
+ gemm1_alpha = quant_config.gemm1_alpha
+ if gemm1_alpha is None:
+ gemm1_alpha = moe_config.swiglu_alpha
+ self.gemm1_alpha = 1.0 if gemm1_alpha is None else gemm1_alpha
+
+ gemm1_beta = quant_config.gemm1_beta
+ if gemm1_beta is None:
+ gemm1_beta = moe_config.swiglu_beta
+ self.gemm1_beta = 0.0 if gemm1_beta is None else gemm1_beta
111 changes: 111 additions & 0 deletions benchmarks/single_node/fixed_seq_len/minimaxm3_fp4_rtx6000pro.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,111 @@
#!/usr/bin/env bash

# MiniMax-M3 NVFP4 RTX PRO 6000 Blackwell single-node vLLM recipe.
# This is the PCIe/SM120 counterpart to minimaxm3_fp4_b200.sh. It keeps
# the ModelOpt NVFP4, FP8 KV-cache, and MSA block-size settings while using
# NCCL collectives instead of the B200-tuned FlashInfer/TRT-LLM all-reduce.
#
# The pinned vLLM image predates the MiniMax-M3 Marlin fixes tracked by
# vLLM PRs #45836 and #48929. Apply the narrow compatibility patch covered
# by docs/waiver/2306.md before importing vLLM.

source "$(dirname "$0")/../../benchmark_lib.sh"

check_env_vars \
MODEL \
TP \
EP_SIZE \
DP_ATTENTION \
CONC \
ISL \
OSL \
MAX_MODEL_LEN \
RANDOM_RANGE_RATIO \
RESULT_FILENAME

if [[ "$MODEL" != /* ]]; then hf download "$MODEL"; fi

if [[ -n "$SLURM_JOB_ID" ]]; then
echo "JOB $SLURM_JOB_ID running on $SLURMD_NODENAME"
fi

nvidia-smi

SERVER_LOG=/workspace/server.log
GPU_MEM_UTIL="${GPU_MEM_UTIL:-0.90}"

export VLLM_ENGINE_READY_TIMEOUT_S=3600
export VLLM_FLOAT32_MATMUL_PRECISION=high

if [ "${DP_ATTENTION}" = "true" ]; then
PARALLEL_ARGS=(
--tensor-parallel-size 1
--data-parallel-size "$TP"
--enable-expert-parallel
)
elif [ "$EP_SIZE" -gt 1 ]; then
PARALLEL_ARGS=(
--tensor-parallel-size "$TP"
--enable-expert-parallel
)
else
PARALLEL_ARGS=(--tensor-parallel-size "$TP")
fi

if [ "${EVAL_ONLY}" = "true" ]; then
setup_eval_context
MAX_MODEL_LEN="$EVAL_MAX_MODEL_LEN"
fi
start_gpu_monitor

VLLM_SITE_PACKAGES="$(
python3 -c 'import sysconfig; print(sysconfig.get_paths()["purelib"])'
)"
VLLM_MARLIN_PATCH=/workspace/benchmarks/patches/vllm/minimax_m3_nvfp4_marlin.patch
if ! patch --batch --forward --fuzz=0 -p1 -d "$VLLM_SITE_PACKAGES" \
< "$VLLM_MARLIN_PATCH"; then
echo "Failed to apply the pinned MiniMax-M3 Marlin compatibility patch" >&2
exit 1
fi
Comment on lines +61 to +69

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 The Marlin compatibility patch step (lines 61-69) hard-aborts if patch --forward returns nonzero, but --forward does not make re-applying an already-applied patch a no-op — GNU patch exits 1 with "Reversed (or previously applied) patch detected!" in that case too. This breaks re-running the script against an already-patched vLLM install (e.g. the repo's own debug-runs fast-iteration workflow of SSHing onto the runner and re-running in the same container), and will also bite if the pinned image ever upstreams just one of the two patch hunks. The sibling minimaxm3_fp8_b200.sh recipe avoids this by explicitly detecting the already-applied state before failing.

Extended reasoning...

The bug: minimaxm3_fp4_rtx6000pro.sh lines 65-69 run:

if ! patch --batch --forward --fuzz=0 -p1 -d "$VLLM_SITE_PACKAGES" < "$VLLM_MARLIN_PATCH"; then
    echo "Failed to apply the pinned MiniMax-M3 Marlin compatibility patch" >&2
    exit 1
fi

The intent of --forward (GNU patch's -N) is presumably to make the script idempotent — i.e., skip hunks that are already applied so re-running the script against a previously-patched vLLM install doesn't fail. That's not what it does. --forward only suppresses applying a patch that would go backwards relative to already-modified source; it does not exit 0 when a hunk is already present. On a second apply of the identical patch, GNU patch prints Reversed (or previously applied) patch detected! Skipping patch., writes a .rej file, and exits 1 — a status this script treats identically to a genuinely broken/mismatched patch.

Reachability: normal CI is unaffected, since runners/launch_rtx6000pro-lat.sh uses docker run --rm, giving a fresh, unpatched container on every invocation. The failure mode is triggered by this repo's own debug-runs skill, which instructs engineers to SSH onto the runner and re-run the benchmark script directly in the same live container (via docker exec/enroot) for fast iteration on things like CONC. The second invocation in that same container hits an already-patched vLLM tree and hard-fails at the patch step before vllm serve even starts, defeating the fast-iteration workflow the skill exists for. It would also resurface once the pinned image upstreams just one of the two patch hunks (per the removal plan in docs/waiver/2306.md) before the patch file is deleted — a partially-applied multi-hunk patch still returns nonzero overall even though the still-needed hunk would apply cleanly.

Step-by-step proof:

  1. First run: container has an unpatched vLLM install. patch --batch --forward --fuzz=0 -p1 -d $VLLM_SITE_PACKAGES < minimax_m3_nvfp4_marlin.patch applies both hunks cleanly, exits 0. vllm serve starts normally.
  2. Engineer SSHes onto the runner mid-debug, re-runs the same script in the same container to tweak CONC.
  3. patch is invoked again against the now-already-patched nvfp4.py and marlin_moe.py. GNU patch detects both hunks are already present, prints "Reversed (or previously applied) patch detected! Skipping patch.", writes .rej, and exits 1.
  4. The script's if ! patch ...; then echo ...; exit 1; fi treats this exit code as failure and aborts before vllm serve is ever invoked — even though the vLLM install is in the exact state the script wants it in.

Why nothing else prevents this: --fuzz=0 and -p1 only control hunk-matching leniency and path stripping; neither affects idempotency. There's no pre-check for an already-patched state, unlike the sibling minimaxm3_fp8_b200.sh recipe (lines ~30-55), which runs a small Python idempotent-patch routine that explicitly prints [minimax-m3-msa-patch] already applied and treats that as success rather than failure — i.e., idempotency is a known, established concern for these runtime-patch recipes that this new recipe doesn't carry over.

Suggested fix: mirror the b200 recipe's approach (an idempotent Python-based patch application, or a pre-check such as patch --dry-run -R to detect 'already applied' and treat it as success rather than failure) so re-running the script against an already-patched install is a no-op instead of a hard failure.

Severity: all three independent verifiers empirically confirmed the behavior with GNU patch and agreed on nit — this is a real, reproducible bug, but normal CI (fresh container per run) is unaffected; it only degrades the manual/debug re-run path and the future partial-upstream scenario, not the sweep this PR is being merged for.


set -x
vllm serve "$MODEL" --port "$PORT" \
"${PARALLEL_ARGS[@]}" \
--disable-custom-all-reduce \
--gpu-memory-utilization "$GPU_MEM_UTIL" \
--max-model-len "$MAX_MODEL_LEN" \
--kv-cache-dtype fp8 \
--block-size 128 \
--language-model-only \
--attention-backend TRITON_ATTN \
--moe-backend marlin \
--max-cudagraph-capture-size 2048 \
--max-num-batched-tokens "$((ISL * 2))" \
--stream-interval 20 \
--no-enable-prefix-caching \
--trust-remote-code > "$SERVER_LOG" 2>&1 &

SERVER_PID=$!

wait_for_server_ready --port "$PORT" --server-log "$SERVER_LOG" --server-pid "$SERVER_PID"

run_benchmark_serving \
--model "$MODEL" \
--port "$PORT" \
--backend vllm \
--input-len "$ISL" \
--output-len "$OSL" \
--random-range-ratio "$RANDOM_RANGE_RATIO" \
--num-prompts "$((CONC * 10))" \
--max-concurrency "$CONC" \
--result-filename "$RESULT_FILENAME" \
--result-dir /workspace/ \
--trust-remote-code

if [ "${RUN_EVAL}" = "true" ]; then
run_eval --framework lm-eval --port "$PORT"
append_lm_eval_summary
fi

stop_gpu_monitor
set +x
19 changes: 19 additions & 0 deletions configs/nvidia-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -7348,6 +7348,25 @@ minimaxm3-fp8-b300-vllm:
- { tp: 4, ep: 4, dp-attn: true, conc-start: 64, conc-end: 128 }
- { tp: 8, ep: 8, dp-attn: true, conc-start: 128, conc-end: 512 }

# MiniMax-M3 NVFP4 single-node vLLM bring-up on 8x RTX PRO 6000.
minimaxm3-fp4-rtx6000pro-vllm:
image: vllm/vllm-openai:vllm-minimax-m3-perf-x86_64-13.0.1-8b00f41@sha256:6af4be7ae69a5f424de85b3d514ac79778bdcf4a9d05f93ec908a4095e0f1253
model: nvidia/MiniMax-M3-NVFP4
model-prefix: minimaxm3
runner: rtx6000pro-lat
precision: fp4
framework: vllm
multinode: false
scenarios:
fixed-seq-len:
- isl: 8192
osl: 1024
search-space:
- { tp: 8, conc-list: [1, 4, 16, 64] }
# Marlin DEP8 uses standard all-gather/reduce-scatter expert exchange;
# DeepEP/NIXL batched activation layouts are not compatible.
- { tp: 8, ep: 8, dp-attn: true, conc-list: [1, 4, 16, 64] }

# MiniMax-M3 NVFP4 (nvidia/MiniMax-M3-NVFP4) B300 single-node vLLM — FP4 variant
# of minimaxm3-fp8-b300-vllm. MiniMax-M3 modelopt NVFP4 support (vllm-project/vllm
# PR #46380) is baked into the perf container image, so no runtime patch is
Expand Down
9 changes: 9 additions & 0 deletions configs/runners.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -164,6 +164,10 @@ labels:
- gb300-nv_0
- gb300-nv_1
- gb300-nv_2
rtx6000pro:
- rtx6000pro-lat_00
rtx6000pro-lat:
- rtx6000pro-lat_00
cluster:h100-cw:
- h100-cw_00
- h100-cw_01
Expand Down Expand Up @@ -251,6 +255,8 @@ labels:
- gb300-nv_0
- gb300-nv_1
- gb300-nv_2
cluster:rtx6000pro-lat:
- rtx6000pro-lat_00
cluster:mi300x-amds:
- mi300x-amds_00
- mi300x-amds_01
Expand Down Expand Up @@ -315,6 +321,9 @@ hardware:
cluster:gb200-nv:
available-cpu-dram-mib: 860_160
gpus-per-node: 4
cluster:rtx6000pro-lat:
available-cpu-dram-mib: 1_500_000
gpus-per-node: 8
cluster:mi300x-amds:
available-cpu-dram-mib: 2_321_924
gpus-per-node: 8
Expand Down
46 changes: 46 additions & 0 deletions docs/waiver/2306.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
# Runtime patch waiver for PR 2306

<div align="center">

**English** | [中文](./2306_zh.md)

</div>

## Scope

The MiniMax-M3 NVFP4 RTX PRO 6000 recipe applies
`benchmarks/patches/vllm/minimax_m3_nvfp4_marlin.patch` to the vLLM Python
package in the pinned container before starting the server. The patch:

- admits Marlin to the clamped NVFP4 MoE backend allowlist; and
- makes Marlin fall back to the model's SwiGLU clamp, alpha, and beta when
those values are absent from the quantization config.

It does not modify a CUDA kernel, checkpoint weight, or benchmark result.

## Why the unmodified image cannot run this benchmark

MiniMax-M3 sets a SwiGLU clamp. The pinned vLLM image consequently filters its
NVFP4 MoE candidates to FlashInfer TRT-LLM, whose device guard accepts SM100
but rejects the RTX PRO 6000's SM120 compute capability. Explicitly selecting
Marlin in the unmodified image is also rejected by that allowlist. If only the
allowlist is widened, Marlin substitutes plain-SiLU defaults for the model's
alpha and beta and produces incorrect output.

The recipe therefore selects `--moe-backend marlin` and carries the minimum
Python adapter fix required to preserve the model's activation parameters.

## Upstream tracking

The Marlin clamp allowlist is upstream in
https://github.com/vllm-project/vllm/pull/45836. The model-config parameter
fallback is the Marlin portion of
https://github.com/vllm-project/vllm/pull/48929. The latter pull request
includes focused tests and reports coherent deterministic MiniMax-M3 NVFP4
output after the fix, where the unpatched Marlin path produced garbled output.

## Removal plan

Replace the pinned image with the first suitable upstream vLLM image that
contains the Marlin fix, then remove the runtime patch application and patch
file. Revalidate TP8 8k/1k generation on SM120 before removing this waiver.
40 changes: 40 additions & 0 deletions docs/waiver/2306_zh.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
# PR 2306 运行时补丁豁免

<div align="center">

[English](./2306.md) | **中文**

</div>

## 范围

MiniMax-M3 NVFP4 RTX PRO 6000 配方会在启动服务前,将
`benchmarks/patches/vllm/minimax_m3_nvfp4_marlin.patch` 应用到锁定镜像中的
vLLM Python 包。该补丁:

- 将 Marlin 加入支持 clamp 的 NVFP4 混合专家(MoE)后端允许列表;以及
- 当量化配置未提供 SwiGLU clamp、alpha 和 beta 时,让 Marlin 回退到模型配置中的对应值。

该补丁不修改 CUDA 内核、模型权重或基准测试结果。

## 为何无法直接使用未修改的镜像

MiniMax-M3 设置了 SwiGLU clamp,因此锁定版本的 vLLM 会把 NVFP4 MoE 候选后端
限制为 FlashInfer TRT-LLM。该后端的设备检查只接受 SM100,不接受 RTX PRO 6000
的 SM120 计算能力。未修改的镜像还会通过允许列表拒绝显式选择 Marlin。如果仅
扩大允许列表,Marlin 会使用普通 SiLU 的 alpha、beta 默认值,并产生错误输出。

因此,该配方显式选择 `--moe-backend marlin`,并携带保留模型激活参数所需的最小
Python 适配层修复。

## 上游跟踪

Marlin 的 clamp 允许列表修复已由
https://github.com/vllm-project/vllm/pull/45836 合入上游。模型配置参数回退逻辑来自
https://github.com/vllm-project/vllm/pull/48929 中的 Marlin 部分。后者包含针对性
测试,并报告修复后 MiniMax-M3 NVFP4 的确定性输出从乱码恢复为连贯文本。

## 移除计划

当包含 Marlin 修复的合适 vLLM 上游镜像发布后,更新锁定镜像,并移除运行时补丁
应用逻辑及补丁文件。移除此豁免前,需要在 SM120 上重新验证 TP8 8k/1k 生成。
13 changes: 13 additions & 0 deletions perf-changelog.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -5060,3 +5060,16 @@
- "Re-pin VLLM_ROUTER_IMAGE to vllm/vllm-router:nightly-20260716-1fbcde7 (previous nightly-20260629-e667ebb was garbage-collected from Docker Hub)"
- "Exclude known-bad nodes mia1-p01-g09,g14 from the disagg node pool"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2301

- config-keys:
- minimaxm3-fp4-rtx6000pro-vllm
scenario-type:
- fixed-seq-len
description:
- "Add MiniMax-M3 NVFP4 single-node vLLM benchmarking on the 8x RTX PRO 6000 Blackwell Latitude runner"
- "Sweep 8k/1k at concurrency 1, 4, 16, and 64 for both GPU-resident TP8 and DEP8 (DP attention with expert parallelism), using FP8 KV cache, MSA block size 128, and NCCL collectives for the PCIe-only topology"
- "Pin the CUDA 13.0.1 MiniMax-M3 performance image by digest and force the non-FlashInfer Marlin MoE backend on SM120"
- "Apply the narrow vLLM adapter fixes tracked by upstream PRs #45836 and #48929 so Marlin preserves MiniMax-M3's SwiGLU clamp, alpha, and beta; documented by docs/waiver/2306.md"
- "Use Triton attention because MiniMax-M3 requires block size 128, which the image's default SM120 FlashInfer attention backend does not accept"
- "Disable NCCL IB/RoCE discovery on this single-node runner to avoid an NCCL 2.28.9 bnxt_re initialization segfault"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2306
Loading
Loading