共 64 个 commit,涉及 303 个文件,+11960/-3245 行变动。

概要

统计项数值
Commit 数64
变更文件303
新增行数+11960
删除行数-3245

Commit 列表

📦 Other

  • 7e85d3a #50126 — [ROCm] Enable pinned memory on supported WSL2 kernels (#50126)
    • 作者: Flora Cui | +29/-3 | 2 个文件

    Assisted-by: GitHub Copilot RocmPlatform did not override is_pin_memory_available(), so ROCm builds running under WSL always fell back to the conservative base Platform.is_pin_memory_available(), which unconditionally disables pinned memory. This adds a RocmPlatform override that mirrors the existing kernel-version gating already used by CudaPlatformBase. While touching this code path, the WSL war…

  • f2bfad9 #50068 — [Model] Enable Qwen3.8 for AMD Rocm (#50068)
    • 作者: haic0 | +10/-0 | 1 个文件
    • register the text-only Qwen3_5ForCausalLM and Qwen3_5MoeForCausalLM architectures - advertise hybrid and M-RoPE support on the causal implementation - expose Gated DeltaNet Mamba cache dtype, shape, and copy metadata so text-only Qwen3.5-compatible checkpoints such as Qwen3.8 Max FP8 can initialize through the causal LM path
  • d5aae2b #51357 — Fix ROCm architecture import on non-ROCm platforms (#51357)
    • 作者: Xiaochang Wu | +20/-9 | 2 个文件
    • guard the ROCm-only on_gfx1250 import with current_platform.is_rocm() - cache the architecture check for AITER backend selection - prevent XPU model loading from initializing torch.cuda through the ROCm platform module ## Testing - pre-commit run –from-ref origin/main –to-ref HEAD - DeepSeek V4 loaded all 46 checkpoint shards and completed generation successfully on TP8 XPU
  • ae934ba #48355 — feat: extended EPLB support for Mistral Large 3 and additional MoE backends (#48355)
    • 作者: Julien Debache | +500/-96 | 16 个文件

    Enable EPLB for MoE models whose quantization config derives per-expert state at load time, and for multi-modal models that nest the MoE language model: 1. Nested MoE models. Add get_mixture_of_experts_model() to resolve the MixtureOfExperts interface through VLM wrappers that don’t implement it themselves. The model runner resolves it once and reuses it. 2. **NVFP4 MoE (compressed-tensors W4A…

  • 8d9b52f #51365 — [XPU] quick fix online quantization UT break (#51365)
    • 作者: Yan Ma | +11/-6 | 1 个文件

    Test Result —

  • 5ec47f3 #50234 — [PD][PushConnector] Record last activity of remotes to allow clean up of stale ones (#50234)
    • 作者: Nicolò Lucchesi | +59/-2 | 3 个文件

    Part of the stability enhancement efforts described here https://github.com/vllm-project/vllm/issues/48633. The PushConnector currently inherits all the logic from base_worker.py necessary to clean up old/stale remotes data structures here and prevent cpu leak (more info here https://github.com/vllm-project/vllm/pull/44424). It is not tapping into that workflow though as the connector itself isn’t…

  • da78833 #47972 — Support DeepSeek-V4 AMD Quark NVFP4 with emulation kernel (#47972)
    • 作者: jimmy-adams | +332/-17 | 11 个文件

    Add support for DeepSeek-V4 AMD Quark mixed-quantized checkpoints, where different parts of the model may use different quantization layouts, including per-block FP8 linear layers and NVFP4 MoE experts. This PR focuses on the Quark/DeepSeek-V4 model-side pieces needed to correctly identify and load these mixed-quantized checkpoints. It avoids changing generic weight-loading behavior and keeps Deep…

  • 21ea5b4 #50902 — [rl] Stateful Trainer Send: NCCL + Sparse NCCL [3/N] (#50902)
    • 作者: Aaron Hao | +1418/-779 | 14 个文件

    Trainer-side weight transfer (3/N): migrate NCCL + sparse NCCL ## Context Third PR of the trainer-side weight-transfer rework. - PR 1 (merged, #48042): introduced the new trainer-side abstractions (WeightSource / ModuleSource, VLLMWeightSyncClient, TrainerWeightTransferEngine, WeightTransferTrainerFactory). Purely additive; no backend migrated. - PR 2 (merged, #48981): migrated the **IPC…

  • 0406ba2 #51341 — fix pre-commit broken (#51341)
    • 作者: Kunshang Ji | +2/-0 | 1 个文件

    Test Result —

  • b1e12d1 #51304 — [V1] Copy NaN-in-logits counts to host asynchronously (#51304)
    • 作者: Nick Hill | +52/-18 | 1 个文件

    The reason for making this change to legacy MRV1 is that we’re planning to turn on nan checking by default for all CI tests, and we don’t want it to effectively disable async scheduling for most of them that (for now) still run using MRV1. VLLM_COMPUTE_NANS_IN_LOGITS made _bookkeeping_sync do logits.isnan().sum(dim=-1).cpu() on the main stream, one blocking D2H per step. That defeats async schedul…

  • 8170c23 #49764 — [Quantization] Share online weight scales across TP (#49764)
    • 作者: Matej Sirovatka | +314/-23 | 5 个文件

    Online weight quantization currently derives scales from rank-local shards. Consequently, tensor-parallel packing can use a different recipe from unsharded or expert-parallel packing of the same BF16 checkpoint. This PR: - reduces the scalar dense-weight amax across the TP group before online FP8 per-tensor packing, only when the linear weight is actually TP-sharded; - reduces the per-expert amax …

  • 27930df #50939 — [Model Runner V2] Fix -1 placeholder draft token ids in rejection sam… (#50939)
    • 作者: Giancarlo Delfin | +229/-74 | 2 个文件

    Summary The block verification path does not currently guard against -1 placeholder draft token ids, which can be used as padding. This PR simply ensures that these draft tokens are rejected, and that they don’t cause OOB for block verification. This was already handled for standard/greedy rejection sampling here: https://github.com/vllm-project/vllm/pull/46533. # Tests Added three tests under t…

  • 9bca7d8 #51300 — docs(governance): refresh committers list, add TSC note, update project leads (#51300)
    • 作者: Simon Mo | +40/-34 | 2 个文件

    This PR refreshes the vLLM governance documentation to reflect the current committer group, formalize the TSC identity, and update the Project Leads list. ### Changes 1. docs/governance/process.md - Clarify that the Core Maintainers / Project Leads committee is also the Technical Steering Committee (TSC) as defined by Linux Foundation Project Governance. - Update Project Leads: remove Lu Fang,…

  • 9464529 #50931 — [ModelRunner v2] Enable decoder token-wise pooling (#50931)
    • 作者: Taneem Ibrahim | +86/-58 | 8 个文件

    Advances #41286 by enabling decoder token-wise pooling on Model Runner V2. The pooling already supports chunked accumulation, but MRV2 filtered token_embed and token_classify for non-encoder models. This removes that flag. This unblocks text decoder embedding models, reward/process-reward models, and rerankers. ## Reproducer Output on Main: Output on this branch: ## Test Plan and Results #…

  • 4d341ca #49206 — fix: resolve silent request skipping in PRIORITY scheduling (#49206)
    • 作者: Tejas | +159/-2 | 2 个文件

    Fixes #49097. This PR resolves a logic bug in the SchedulingPolicy.PRIORITY preemption path. When a request is preempted, the scheduler’s request bookkeeping (req_index) was not correctly adjusted if the preempted request appeared earlier in the self.running list than the current iteration cursor. This caused the subsequent request to be silently skipped for the entire scheduling step. 1. Created …

  • c5d470a #51210 — [ModelRunner V2] Minor indexing optimizations (#51210)
    • 作者: Nick Hill | +6/-6 | 3 个文件
    • Use getitem rather than get with np.fromiter(map(…)) (~6% faster) - Use np.intp for indexing types (avoids copy/type-conversion every time it’s used) - Use torch.int64 for gpu tensor index mapping (same reason) Idea came from this PR: https://github.com/vllm-project/vllm/pull/50532
  • adc3e03 #48977 — [Mypy Fix] Mypy fix for “vllm/model_executor/models/[aA][bB]” (#48977)
    • 作者: Wentao Ye | +214/-167 | 23 个文件

    Part of https://github.com/vllm-project/vllm/issues/26533 Originally Now

  • d8eabdb #50578 — [ROCm][MLA] Use asm decode for non-divisor small head counts (#50578)
    • 作者: vanshbhatia-amd | +393/-25 | 4 个文件

    On ROCM_AITER_MLA, MLA decode with fewer than 16 query heads per rank runs the Gluon decode kernel (AiterMLAHelper.use_gluon_decode). Gluon parallelizes only over heads, so a single workgroup marches through the whole KV cache and per-token latency scales linearly with context length. The faster asm persistent decode requires num_heads >= 16; smaller head counts are padded up to 16, but get_mla_…

  • 81be2e0 #49601 — [Weight processing] Copy over new_data attributes in replace_parameter (#49601)
    • 作者: fxmarty-amd | +334/-51 | 10 个文件

    replace_parameter re-wraps new_data into a fresh torch.nn.Parameter, which silently drops any custom Python attribute previously attached to the tensor (e.g. kernel dispatch flags such as is_shuffled). Several process_weights_after_loading implementations worked around this by re-setting such attributes on layer. right after each replace_parameter call, or, in the AITER MXFP8 MoE experts ca…

  • b38e111 #50613 — [Attention][MLA] Per-request scheduling for MLA chunked context (#50613)
    • 作者: Matthew Bonanni | +940/-552 | 15 个文件

    Implements #50497. MLA prefill chunks are fit into the available workspace rather than forced to be the same size. This can reduce the overall number of chunks and improve prefill latency. ## Validation ### Correctness pytest tests/v1/attention/test_mla_context_chunks.py -q passes GSM8K with DeepSeek-V2-Lite-Chat, TP=4, DCP=2, FlashMLA, prefix caching: | Revision | Strict exact match | Flexible ex…

  • e7b8d59 #51247 — Fully generalise input embedding handling in Transformers modelling backend (#51247)
    • 作者: Harry Mellor | +300/-47 | 6 个文件

    Before this PR we replaced the entire module returned by get_input_embeddings with VocabParallelEmbedding and had a special case for scaled input embeddings. This does not generalise well, particularly if input embeddings perform additional operations that we have not accounted for. This PR adds replace_embedding_class which: - Fully replaces with VocabParallelEmbedding if the embedding was a bare…

  • 7b4ed49 #50185 — attn_res kernel latency improvements (#50185)
    • 作者: gnovack | +84/-39 | 1 个文件

    A few small improvements to reduce the latency of the attention residual kernel: - Use vectorized loads when populating q_cache - Treat B and N as compile-time constants - Set NC = 3 for low batch size path - Unroll the main loop over num_chunks ## Microbenchmark result Microbenchmarks run on GB300, stacking each of the above changes Tokens | Baseline | q_cache vector load | template B & N | NC=3 …

  • 41e7746 #49599 — Update vllm to point to flash-attention commit that builds FA3 with torch stable API. (Retry) (#49599)
    • 作者: Chris Leonard | +2/-5 | 2 个文件

    Re-attempt of FA3 stable ABI migration like the one in #46644 Points to the top commit on the PR https://github.com/vllm-project/flash-attention/pull/165. build with new flash-attention commit and check that only stable symbols are exposed. ## Test Result build succeeded and library is torch abi stable (see below) — Migration progress of vLLM using the Audit Python extension torch-abi-audit:

  • 62a8631 #51224 — [VocabParallelEmbedding] fix extra_repr fields concat (#51224)
    • 作者: Ning Xie | +2/-1 | 1 个文件

    Chore fix. correct VocabParallelEmbedding extra_repr fields with corresponding num_embeddings and num_embeddings_per_partition NA ## Test Result NA —

  • 1e05b21 #51242 — Remove the XPU branch of topk_softplus_sqrt (#51242)
    • 作者: Liangqiusong | +0/-62 | 1 个文件

    The topk_softplus_sqrt operator has now been merged into vllm_xpu_kernels, so we can remove the XPU Torch branch.

  • 865781e #50289 — [Rust Frontend] Add standalone Rust renderer (#50289)
    • 作者: Sage | +720/-40 | 15 个文件

    This implements Stage 2 of vllm-project/vllm#49047 after the Stage 1 Adds an engine-free vllm-rs render command that reuses the Rust frontend’s existing chat and text request-preparation pipeline without starting or connecting to an inference engine. The server exposes: - GET /health - GET|POST /ping - GET /v1/models - POST /v1/chat/completions/render - POST /v1/completions/render The implementati…

  • 22013f7 #51149 — Interns2mobius support (#51149)
    • 作者: Lyu Han | +664/-3 | 8 个文件

    Add support for model https://huggingface.co/internlm/Intern-S2-Mobius ## Test Result —

  • 2fa4904 #50411 — [Model] Fused mm preprocess normalisation on the Device (#50411)
    • 作者: wang.yuqi | +359/-11 | 13 个文件

    This is one of my favorite CV tricks. 1. When normalize is computed on the GPU, it can be fused into subsequent operators, or even if left unfused, it takes almost no time. 2. Currently, image data transmission from the entrypoint to the engine core and then to the GPU uses uint8, rather than the model dtype bf16, and is only converted to bf16 on the GPU before entering input normal. Moreover, thi…

  • 9c22668 #50029 — [Quantization] Preserve precision in online NVFP4 expert packing (#50029)
    • 作者: Matej Sirovatka | +58/-10 | 2 个文件

    Online NVFP4 MoE packing currently folds each expert’s FP32 global encode scale into its BF16/FP16 weight, casts the scaled tensor back to the original dtype, and then quantizes with a neutral scale. That intermediate cast adds a rounding step before the group-16 scales and E2M1 values are selected. This change passes each original expert tensor and its FP32 global encode scale directly to scaled_…

  • 276f0bb #50946 — [XPU] Register fake meta kernel for fp4_gemm (#50946)
    • 作者: Chaojun Zhang | +18/-0 | 1 个文件

    Issue: torch.ops._xpu_C.fp4_gemm lacked a fake kernel, causing graph breaks under torch.compile (default) and disrupting downstream compilation. ## Fix: Added _fp4_gemm_fake: output shape (M, N) with M = input.flatten(0,-2).size(0), N = weight.size(-1), dtype from out_dtype or torch.get_default_dtype(). This enables shape propagation and keeps the graph intact.

  • 13726c8 #49453 — [CPU] Add MLA backend so DeepSeek-V2/V3 can run on CPU (#49453)
    • 作者: maobaolong | +367/-14 | 6 个文件

    Make DeepSeek-V2/V3 style MLA models runnable on the CPU platform. This is a reference-quality path (correctness first, performance later) so people without a GPU handy can at least kick the tires on MLA models locally. The plumbing is: - New CPUMLABackend / CPUMLAImpl under vllm/v1/attention/backends/mla/. It inherits MLACommonBackend / MLACommonImpl so the shared MLA scaffolding (weight-absorbed…

  • ad52802 #50904 — [GLM Perf] DSv32/glm use skip topk for MTP case, 2.0x kernel performance improvement (#50904)
    • 作者: Wentao Ye | +7/-6 | 2 个文件

    Part of https://github.com/vllm-project/vllm/issues/46654 Now we can use skip_topk attribute to skip some redundant calculation In main: Now Acc covered in current unit test Perf can be seen in this AI generated script And we can get

  • febea17 #50355 — [Model] Fix weight prefix mapping for native Qwen3.5 text-only checkp… (#50355)
    • 作者: zofia | +13/-2 | 2 个文件

    Native Qwen3.5 text-only checkpoints (e.g. imdatta0/small_qwen3_5_20b, Timersofc/Qwen3.5-Creative-19B-A3B-REAP) store weights with model.language_model.* prefix, inherited from the VL training framework. The text-only Qwen3_5ForCausalLMBase does not have a language_model submodule, causing AutoWeightsLoader to fail with: ValueError: There is no module or parameter named language_model Add a Weight…

  • 470297c #50980 — [HPC Attention Backend] hpc attention backend support bf16 kv cache with fp8 weight (#50980)
    • 作者: Cheng Jiang | +75/-36 | 3 个文件

    At previous PR #46020 , hpc attention backend (–attention_backend HPC_ATTN) only support fp8 kv cache (–kv-cache-dtype fp8_e4m3) on fp8 model, for example Hy3-FP8. This PR add bfloat16 kv cache (–kv-cache-dtype bfloat16 or –kv-cache-dtype auto) support. You can run FP8 model (not only Hy3-FP8) with bfloat16 kv cache dtype. Hy3-FP8: Qwen3-30B-A3B-Instruct-2507-FP8: ## Test Result —

🐛 Bug Fix

  • 448344c #50965 — [Bugfix] Fix get_open_port() livelock on DP-reserved ports and cover get_open_ports_list (#50965)
    • 作者: aoshen02 | +115/-8 | 2 个文件

    Fixes #50024. get_open_port() must return a port that is bindable and outside the range reserved for the data parallel master (VLLM_DP_MASTER_PORT, +10)). The old implementation rejected a reserved candidate and retried without advancing any state. With VLLM_PORT set the scan is deterministic — same start, same answer — so a free-but-reserved candidate (e.g. VLLM_PORT == VLLM_DP_MASTER_POR…

  • 47228db #48534 — [Bugfix][KV-transfer] MoRIIO: per-layer READ-completion barrier in wait_for_layer_load (#48534)
    • 作者: limeward | +86/-13 | 2 个文件

    In MoRIIO READ mode the decode worker posts a request’s per-layer RDMA reads asynchronously in start_load_kv. Each read completes later when a CQ-poll thread flips its status to Succeeded. wait_for_layer_load() was a no-op (pass), so the attention kernel for a layer could run before that layer’s remote KV had actually landed. Under a full CUDA graph the missing barrier is silently skip…

  • d4ecb75 #48758 — [PD][NixlPush][Bugfix] Fix prefix caching (#48758)
    • 作者: Nicolò Lucchesi | +189/-73 | 5 个文件

    Addressing C3 here https://github.com/vllm-project/vllm/issues/48633. On this mode, P needs to send the full sequence to D, regardless of local prefix cache hits; on the other hand, D must take local hits into account, and only request blocks that does not have locally (uncomputed). This means that on P we need to tail-trim and only WRITE/send the last N missing/uncomputed blocks. (side-note: …

  • a231c5c #51260 — [Bugfix] Skip fetching revision for model when model and weights_model are different (#51260)
    • 作者: music-dino | +27/-7 | 2 个文件

    After PR #49990, two gguf tests soft failed on the nightly https://buildkite.com/vllm/ci/builds/82629/list?sid=019fd5a9-db3f-42ba-9b21-993c983b2c7d&tab=output - plugins_tests/gguf/test_gguf_plugin_generate.py::test_models[1-8-32-bfloat16-model0] - plugins_tests/gguf/test_gguf_plugin_generate.py::test_models[1-8-32-bfloat16-model1] with the following error: ### Root cause ModelConfig can load confi…

  • 4f76c8a #50393 — [Bugfix][Platform] Stop re-initializing NVML on every device-capability check (fixes #50381) (#50393)
    • 作者: Sebastian Woo | +82/-1 | 2 个文件

    Fixes #50381. NvmlCudaPlatform.has_device_capability (which CudaPlatform resolves to wherever NVML is available) paid a full nvmlInit()/nvmlShutdown() pair on every call while doing nothing that required an NVML context. ## Root cause The two sibling methods in vllm/platforms/cuda.py were decorated inconsistently: with_nvml_context is just nvmlInit() → call → nvmlShutdown(). The body of has_device…

  • b8db7f4 #50833 — [Bugfix][Quantization] Fix dynamic INT8 W8A8 MoE config being built as W8A16 (#50833)
    • 作者: Hank_ | +5/-4 | 1 个文件

    Fix the runtime quantization config for compressed-tensors INT8 W8A8 MoE models with dynamically quantized per-token activations. TL;DR: For the Triton and CPU INT8 MoE backends, this configuration is currently misclassified as W8A16 while constructing FusedMoEQuantConfig. As a result, the per-token INT8 activation quantization requested by the checkpoint is disabled. For int8_w8a8 compres…

  • c810e5e #51100 — [Bugfix] Fix Mamba all-mode CPU offload boundary alignment (#51100)
    • 作者: Qianxu Wang | +21/-13 | 2 个文件

    Fixes #51094. - Treat mamba_cache_mode=“all” like “align” when constraining OffloadingConnector hit windows. - Recompute the last token at an exact offload-chunk boundary instead of restoring a Mamba point state that already includes it. - Extend the existing Mamba CPU-offload boundary test to cover both cache modes. ## Root cause resolve_mamba_align_size() only enabled boundary alignment for mamb…

  • dd856e4 #51222 — [Bugfix][EPD][Model Runner V2] Skip gather mm embeddings for encoder only instance (#51222)
    • 作者: Tianyu Guo | +101/-11 | 5 个文件

    An EPD encoder instance dies with RuntimeError: Encoder cache miss as soon as it is handed a multi-modal item the EC connector already holds. The whole instance goes down, not just the request: everything after it gets a connection error. Root cause. When ec_connector.has_cache_item() is true the scheduler stops scheduling that item for encoding (scheduler.py:1610) — correct, the embedding is …

  • d6af803 #50276 — [Bugfix] Fix packed KV block zeroing stride (#50276)
    • 作者: wangxian001 | +81/-17 | 2 个文件

    Follow up on #49704, which added support for non-uniform page sizes in KVBlockZeroer. KVBlockZeroer still uses a segment’s page size for both the block-to-block address stride and the number of elements to clear. For a packed KV view, the physical distance between consecutive logical blocks can be larger than that segment’s page. This can clear adjacent packed data, and clearing the final logical …

  • c56f169 #51113 — [Bugfix] Keep mamba align prefill chunks block-aligned past last_cache_position (#51113)
    • 作者: Yifan Qiao | +258/-11 | 2 个文件

    Credit > > The original work here is @kodek’s (#45477) and @yanghui1-arch’s (#47861). They found this root cause and wrote this fix first; both are co-authors of this PR. > > Their branches went stale on rebase, not on review. What I added is the same fix re-derived against current main, the regression tests, and the measurements below. Fix prefix-cache poisoning on hybrid Mamba/GDN mode…

  • 46e6a83 #51227 — [Bugfix][KV Offload] Clean up resources after initialization failure (#51227)
    • 作者: AlexHuang | +85/-53 | 2 个文件

    [Bugfix][KV Offload] Clean up resources after initialization failure Prevent /dev/shm/vllm_offload_*.mmap and tier resources from leaking when KV offloading initialization fails. - Clean up the scheduler mmap when primary-tier construction fails. - Shutdown already-created secondary tiers in reverse order when a later tier fails, then continue with primary cleanup. - Clean up CPU/tiering worker …

  • 5fba75a #51249 — [Bugfix][Model] Add missing fused_qkv_a_proj to Kimi-Linear packed_modules_mapping (#51249)
    • 作者: yjz | +1/-0 | 1 个文件

    While testing a Kimi-K3 W4A8 quantized checkpoint, we noticed that several Linear layers in the checkpoint are intentionally left unquantized (bf16) rather than packed as W4A8 — attention projections among them. Digging into why one of those bf16 layers was being treated as quantized led to this bug. KimiLinearModel.packed_modules_mapping (added in #50500) lists the fused params for MoE/dense MLP …

  • ef2615c #51116 — [Bugfix][KV Offload] Fall back when MADV_POPULATE_WRITE is unsupported (#51116)
    • 作者: AlexHuang | +197/-6 | 2 个文件

    Fixes #44888 ## Why this PR SharedOffloadRegion.init unconditionally issues madvise(MADV_POPULATE_WRITE, …) on its shared mmap. On Linux < 5.14 (RHEL 8 ships 4.18) the syscall returns EINVAL, so any vLLM startup that touches TieringOffloadingSpec aborts before KV offloading can register. This PR probes the advice once after mmap creation and falls back to a no-mutation NumPy read-modify-write …

  • 872fd59 #51002 — [Bugfix][LoRA] Guard TrtLlm BF16 MoE LoRA gate on activation type (#51002)
    • 作者: anhtra3889 | +10/-12 | 1 个文件

    Fix BF16 MoE + LoRA backend selection for non-gated models (e.g. NemotronH). The LoRA path in select_unquantized_moe_backend() returns early via _trtllm_bf16_lora_supported() before the modular-kernel oracle runs, so activation / act_and_mul checks on TrtLlmBf16LoRAExperts were skipped. Non-gated experts were routed into the gated-only FlashInfer kernel and crashed at startup during profile_run wi…

🔩 Misc

  • 6b5bec7 #45694 — [Misc] Add and enable Triton kernel unit tests on XPU (#45694)
    • 作者: pmanczak | +58/-13 | 5 个文件

    Makes five Triton kernel unit tests run on Intel GPU (XPU) as well as CUDA, and adds coverage for the round_int8 kernel. Part of RFC #48480. Test-only; no kernel or production code touched. ## Test Result Tested on Arc Pro B70, torch 2.12.0+xpu, triton-xpu 3.7.1 and on H200 on B70: | file | passed | | — | — | | test_block_int8.py | 32 | | test_int8_kernel.py | 44 | | test_triton_scaled_mm.py |…

  • a07086e #50827 — [Misc] Upgrade fastsafetensors version, fix metadata is null (#50827)
    • 作者: rongfu.leng | +9/-9 | 9 个文件

    When use –load-format fastsafetensors this mode MiniMaxAI/MiniMax-M3-MXFP8, it will faild, error info is SafeTensorsMetadataKeyError: ‘data_offsets’ About PR: https://github.com/foundation-model-stack/fastsafetensors/pull/89 ## Test Result —

  • e07532b #50981 — [MISC][Bench] refactor throughput and reuse serve’s get samples (#50981)
    • 作者: Jared Wen | +228/-231 | 3 个文件

    #50849 #50838 vllm bench serve’s datasets is the superset of vllm bench throughput let throughput reuse the get_samples like serve, we can make these two share the same datasets without manually adding each dataset to throughput or serve and we can eliminate the tons of if else in throughput for maintenance ## Test Result —

🔧 Refactor

  • c84789c #51051 — [Refactor] Remove kernel dead code (#51051)
    • 作者: Wentao Ye | +0/-396 | 8 个文件

    Remove kernel dead code

🧪 CI/Tests

  • 0de0362 #48847 — [ROCm][CI] Loosen block-FP8 fused MoE test tolerance for large-K shapes (#48847)
    • 作者: stefankoncarevic | +56/-11 | 2 个文件

    tests/kernels/moe/test_block_fp8.py::test_w8a8_block_fp8_fused_moe compares two Triton block-FP8 fused-MoE kernels (fused_experts and modular_triton_fused_moe) against a native torch reference. On ROCm/gfx950 (MI355X) the large-K (K=7168), large-N (N >= 1024) DeepSeek-style shapes failed the comparison at the base tolerance. Digging in, the divergence turned out to be **reference artifacts, not ke…

  • b706fd1 #51337 — [CI][XPU] Work around intermittent segfault in Intel XPU CI with VLLM_DISABLE_COMPILE_CACHE=1 (#51337)
    • 作者: Chaojun Zhang | +10/-5 | 5 个文件

    Certain Intel XPU CI jobs set VLLM_DISABLE_COMPILE_CACHE=1 to mitigate an intermittent segfault in the Intel Triton XPU backend. Upstream issue: https://github.com/intel/intel-xpu-backend-for-triton/issues/7682

  • 5ac2684 #51293 — [CI] Re-enable FI autotune in GSM8K config for Qwen3.5-35B-A3B (#51293)
    • 作者: Artem Perevedentsev | +0/-1 | 1 个文件

    Recently [Bug]: GPU coredump during FlashInfer trtllm_bf16_moe autotune with Qwen3.5-35B-A3B on B200 (DP=2 + EP) #46083 found bug in FI MoE kernel autotuning. Because of this bug CI job LM Eval Qwen3.5 Models (B200) was failing. To fix CI job’s failure [[Bugfix] fix qwen3.5 ep weight loading #45002](https://github.com/vllm-project/vllm/pull/4500…

  • d35eb6c #51046 — [CI] Exclude KV-connector subtree from broad source dependencies (#51046)
    • 作者: Nicolò Lucchesi | +59/-0 | 13 个文件

    Steps that declare the broad vllm/ or vllm/distributed/ source_file_dependencies run whenever any file under those trees changes. The KV connectors live at vllm/distributed/kv_transfer/, so a self-contained change there — already covered by the dedicated Disaggregated jobs — fans out to the entire model / multimodal / entrypoint / basic-correctness matrix. Concretely, the small nixl push-c…

  • 4f851be #51273 — [ROCm][CI] Update AITER AR+RMS e2e fusion counts for final-norm coverage (#51273)
    • 作者: Divakar Verma | +2/-2 | 1 个文件

    PR #50802 added AiterAllreduceFusedAddRMSNormOutputOnlyPattern, which lets the AITER AllReduce+RMSNorm fusion pass also fuse the terminal model.norm (the fused_add_rms_norm whose residual output is dead). That PR updated the unit test but not the e2e expected counts in tests/compile/fusions_e2e/models.py, so the AITER e2e count is now off by one. This bumps aiter_ar_rms_fusion from n_layers * 2 to…

  • 566c80e #51271 — [CI] Run basic fullgraph correctness on one GPU (#51271)
    • 作者: Michael Goin | +0/-22 | 2 个文件
    • remove test_basic_correctness.py from the 2-GPU and 4-GPU distributed Buildkite steps - remove the remaining PP=2, TP=2 parameter from that file - keep the Granite and embedding coverage in the 1-GPU H200 fullgraph suite ## Why PR #51074 replaced the only 2-GPU model in this test with a TP=1 Granite model. The test’s permissive GPU-count check then caused the TP=1 cases to run redundantly in the…

✨ New Feature

  • 58fcaa0 #49644 — [Feat][Core] Add disk offloading support to SimpleCPUOffloadConnector (#49644)
    • 作者: Guanyi Chen | +466/-19 | 4 个文件

    Implements the disk backend extension envisioned in RFC #19854 (pluggable offloading backends). While #45036 provides SSD offload via the external Mooncake Store connector, this PR adds a native vLLM disk offloading path with zero external dependencies — targeting scenarios where host DRAM is limited but local NVMe capacity is abundant (e.g., dense inference nodes with most memory reserved for…

  • 72c0d67 #46727 — [Feat] Support thinking_token_budget in Model Runner V2 (#46727)
    • 作者: Chauncey | +1122/-43 | 10 个文件

    In some quantized models, such as GLM-5.2 or Qwen quantized models, the model may generate long reasoning traces. Support thinking_token_budget in Model Runner V2 see e2e ## Test Result — BEFORE SUBMITTING, PLEASE READ https://docs.vllm.ai/en/latest/contributing (anything written below this line will be removed by GitHub Actions)

⚡ Performance

  • 43d691e #51253 — [ROCm][Perf] Kimi-K3 Shard Latent MoE up-projection for ROCm path (#51253)
    • 作者: kliuae | +453/-0 | 4 个文件

    For Kimi-K3, the default routed up-projection is a ReplicatedLinear where every TP rank holds the same weight and computes idential full projection where only 1/8 of the work is needed. On the CUDA path, there is LatentMoERunner that shards the workload in a column parallel manner. However, currently the ROCm path lacks this and does not pass a runner_cls to the FusedMoeFactory so it never reaches…

  • 2dfb8ba #50230 — [Perf][CUDA] Programmatic dependent launch for the DSA decode kernels (#50230)
    • 作者: xiaozhoupy | +54/-10 | 2 个文件

    What this does Chain the DSA decode kernels with programmatic dependent launch so back-to-back small kernels overlap their launch latency: - the fused norm/rope and fused-q Triton kernels get a USE_PDL constexpr with gdc_wait() / gdc_launch_dependents(), launched with launch_pdl - the NVFP4 quant kernels switch to cudaLaunchKernelEx with programmatic stream serialization, guarded on __CUDA_ARCH…

🖥️ Kernel

  • e081112 #47106 — [Kernel] Support Nvfp4 Cutedsl Moe Swiglu-oai and Relu2(non-gated) Activation (#47106)
    • 作者: Bi Tiekai | +216/-14 | 9 个文件

    FlashInfer modified the CuteDSL NVFP4 MoE kernel to support the SwiGLU-OAI activation (https://github.com/flashinfer-ai/flashinfer/pull/3737). The CuteDSL NVFP4 MoE kernel now also supports ReLU² non-gated activation. This PR achieves compatibility with both new activations by passing the activation type and related parameters into the CuteDSL MoE kernel. Test input layout prepare: .venv/bin/pytho…

🦀 Rust Frontend

  • 7b9f2da #43417 — [Frontend] Watch frontend processes during engine startup (#43417)
    • 作者: Bugen Zhao | +152/-42 | 5 个文件

    Fail promptly when a frontend process exits while engine cores are still initializing. launch_core_engines yields control so the caller can start API server processes before the engine startup barrier runs. The barrier previously watched local engine cores and the DP coordinator only. A Python API server or Rust frontend that exited during this window could leave the parent waiting for engine star…