共 77 个 commit,涉及 304 个文件,+14386/-2105 行变动。
概要
| 统计项 | 数值 |
|---|---|
| Commit 数 | 77 |
| 变更文件 | 304 |
| 新增行数 | +14386 |
| 删除行数 | -2105 |
Commit 列表
📦 Other
- 7b4ed49 #50185 — attn_res kernel latency improvements (#50185)
- 作者: gnovack | +84/-39 | 1 个文件
A few small improvements to reduce the latency of the attention residual kernel: - Use vectorized loads when populating q_cache - Treat B and N as compile-time constants - Set NC = 3 for low batch size path - Unroll the main loop over num_chunks ## Microbenchmark result Microbenchmarks run on GB300, stacking each of the above changes Tokens | Baseline | q_cache vector load | template B & N | NC=3 …
- 41e7746 #49599 — Update vllm to point to flash-attention commit that builds FA3 with torch stable API. (Retry) (#49599)
- 作者: Chris Leonard | +2/-5 | 2 个文件
Re-attempt of FA3 stable ABI migration like the one in #46644 Points to the top commit on the PR https://github.com/vllm-project/flash-attention/pull/165. build with new flash-attention commit and check that only stable symbols are exposed. ## Test Result build succeeded and library is torch abi stable (see below) — Migration progress of vLLM using the Audit Python extension torch-abi-audit:
- 62a8631 #51224 — [VocabParallelEmbedding] fix extra_repr fields concat (#51224)
- 作者: Ning Xie | +2/-1 | 1 个文件
Chore fix. correct VocabParallelEmbedding extra_repr fields with corresponding num_embeddings and num_embeddings_per_partition NA ## Test Result NA —
- 1e05b21 #51242 — Remove the XPU branch of topk_softplus_sqrt (#51242)
- 作者: Liangqiusong | +0/-62 | 1 个文件
The topk_softplus_sqrt operator has now been merged into vllm_xpu_kernels, so we can remove the XPU Torch branch.
- 865781e #50289 — [Rust Frontend] Add standalone Rust renderer (#50289)
- 作者: Sage | +720/-40 | 15 个文件
This implements Stage 2 of vllm-project/vllm#49047 after the Stage 1 Adds an engine-free vllm-rs render command that reuses the Rust frontend’s existing chat and text request-preparation pipeline without starting or connecting to an inference engine. The server exposes: - GET /health - GET|POST /ping - GET /v1/models - POST /v1/chat/completions/render - POST /v1/completions/render The implementati…
- 22013f7 #51149 — Interns2mobius support (#51149)
- 作者: Lyu Han | +664/-3 | 8 个文件
Add support for model https://huggingface.co/internlm/Intern-S2-Mobius ## Test Result —
- 2fa4904 #50411 — [Model] Fused mm preprocess normalisation on the Device (#50411)
- 作者: wang.yuqi | +359/-11 | 13 个文件
This is one of my favorite CV tricks. 1. When normalize is computed on the GPU, it can be fused into subsequent operators, or even if left unfused, it takes almost no time. 2. Currently, image data transmission from the entrypoint to the engine core and then to the GPU uses uint8, rather than the model dtype bf16, and is only converted to bf16 on the GPU before entering input normal. Moreover, thi…
- 9c22668 #50029 — [Quantization] Preserve precision in online NVFP4 expert packing (#50029)
- 作者: Matej Sirovatka | +58/-10 | 2 个文件
Online NVFP4 MoE packing currently folds each expert’s FP32 global encode scale into its BF16/FP16 weight, casts the scaled tensor back to the original dtype, and then quantizes with a neutral scale. That intermediate cast adds a rounding step before the group-16 scales and E2M1 values are selected. This change passes each original expert tensor and its FP32 global encode scale directly to scaled_…
- 276f0bb #50946 — [XPU] Register fake meta kernel for fp4_gemm (#50946)
- 作者: Chaojun Zhang | +18/-0 | 1 个文件
Issue: torch.ops._xpu_C.fp4_gemm lacked a fake kernel, causing graph breaks under torch.compile (default) and disrupting downstream compilation. ## Fix: Added _fp4_gemm_fake: output shape (M, N) with M = input.flatten(0,-2).size(0), N = weight.size(-1), dtype from out_dtype or torch.get_default_dtype(). This enables shape propagation and keeps the graph intact.
- 13726c8 #49453 — [CPU] Add MLA backend so DeepSeek-V2/V3 can run on CPU (#49453)
- 作者: maobaolong | +367/-14 | 6 个文件
Make DeepSeek-V2/V3 style MLA models runnable on the CPU platform. This is a reference-quality path (correctness first, performance later) so people without a GPU handy can at least kick the tires on MLA models locally. The plumbing is: - New CPUMLABackend / CPUMLAImpl under vllm/v1/attention/backends/mla/. It inherits MLACommonBackend / MLACommonImpl so the shared MLA scaffolding (weight-absorbed…
- ad52802 #50904 — [GLM Perf] DSv32/glm use skip topk for MTP case, 2.0x kernel performance improvement (#50904)
- 作者: Wentao Ye | +7/-6 | 2 个文件
Part of https://github.com/vllm-project/vllm/issues/46654 Now we can use skip_topk attribute to skip some redundant calculation In main: Now Acc covered in current unit test Perf can be seen in this AI generated script And we can get
- febea17 #50355 — [Model] Fix weight prefix mapping for native Qwen3.5 text-only checkp… (#50355)
- 作者: zofia | +13/-2 | 2 个文件
Native Qwen3.5 text-only checkpoints (e.g. imdatta0/small_qwen3_5_20b, Timersofc/Qwen3.5-Creative-19B-A3B-REAP) store weights with model.language_model.* prefix, inherited from the VL training framework. The text-only Qwen3_5ForCausalLMBase does not have a language_model submodule, causing AutoWeightsLoader to fail with: ValueError: There is no module or parameter named language_model Add a Weight…
- 470297c #50980 — [HPC Attention Backend] hpc attention backend support bf16 kv cache with fp8 weight (#50980)
- 作者: Cheng Jiang | +75/-36 | 3 个文件
At previous PR #46020 , hpc attention backend (–attention_backend HPC_ATTN) only support fp8 kv cache (–kv-cache-dtype fp8_e4m3) on fp8 model, for example Hy3-FP8. This PR add bfloat16 kv cache (–kv-cache-dtype bfloat16 or –kv-cache-dtype auto) support. You can run FP8 model (not only Hy3-FP8) with bfloat16 kv cache dtype. Hy3-FP8: Qwen3-30B-A3B-Instruct-2507-FP8: ## Test Result —
- 7e724dc #50802 — [ROCm] Fix AITER all-reduce fusion coverage (#50802)
- 作者: Andreas Karatzas | +95/-46 | 2 个文件
- Register an output-only AITER add-RMSNorm fusion pattern for graphs where the residual result is dead. - Keep the existing two-output pattern unchanged for callers that consume both RMSNorm and residual outputs. - Restructure the grouped-FP8 test model so four independent fusion sites do not compound quantization error through a lossy chain. - Validate the two normal and two indexer fusion varia…
- 8c0f029 #50236 — [XPU] update warning of XPU Graph (#50236)
- 作者: liuzhenwei | +31/-7 | 2 个文件
After upgrading to oneAPI 2026.0 in https://github.com/vllm-project/vllm/pull/48677, the earlier issues with memory and Flash Attention only supporting piecewise mode have been resolved. XPU Graph still only supports single-GPU execution, this PR updates the warning message. ## Test Result —
- b976e98 #50840 — [XPU] Route AWQ linear through choose_mp_linear_kernel (#50840)
- 作者: zofia | +11/-71 | 2 个文件
Purpose Route XPU AWQ linear through choose_mp_linear_kernel (selecting XPUwNa16LinearKernel) instead of the bespoke AutoAWQXPULinearMethod, matching CPU/CUDA. XPU now returns AutoAWQMarlinLinearMethod, which converts AWQ to the standard GPTQ format before handing off to the kernel. This also prevents XPU from falling through to AutoAWQLinearMethod, whose apply() calls the CUDA-only ops.awq_de…
- d779835 #48476 — [XPU] Support MXFP8 linear weights for INC DeepSeek V4 model (#48476)
- 作者: Xiaochang Wu | +36/-5 | 1 个文件
Enable the DeepSeek V4 XPU output-projection path to consume compressed-tensors MXFP8 weights efficiently. - Preprocess weight scale to match oneDNN expected format - For BMM layers such as DeepSeek V4 wo_a, precompute contiguous bmm_weight [G, K, N] and bmm_scale [G, K/32, N/G]. ## Testing Model: https://huggingface.co/INCModel/DeepSeek-V4-Flash-MXFP4-Mixed-CT-AutoRound GSM8K on 8x B70: 0.948. ##…
- 777b01d #45254 — [MM][CG] Support ViT full CUDA graph for Ernie-4.5-VL image inference (#45254)
- 作者: Qiuyang Yue | +289/-13 | 4 个文件
Adds encoder (ViT) CUDA graph support for Ernie4_5_VLMoeForConditionalGeneration (image inputs), under the ViT Full CUDA Graph tracker #38175. Follows the Qwen3-VL SupportsEncoderCudaGraph reference. ## Key Changes - Splits the inline rotary / cu_seqlens / max_seqlen computation out of Ernie4_5_VisionTransformer.forward into prepare_encoder_metadata(), so it can be precomputed on the host and fed …
- 2e09247 #51215 — [Docs] List Intel XPU attention backends (#51215)
- 作者: baodi | +1/-0 | 1 个文件
Document the attention backends available on Intel XPU platforms in the quickstart guide. ## Modifications - Added the supported Intel XPU attention backends: FLASH_ATTN, TRITON_ATTN, TRITON_MLA, XPU_MLA_SPARSE, TORCH_SDPA, and TURBOQUANT. ## Testing Not run; documentation-only change. ## Additional context This change does not affect runtime behavior or model outputs and does not duplicate an exi…
- d54b58c #51107 — [Hardware] Use torch.accelerator.empty_host_cache() for host cache cl… (#51107)
- 作者: liuzhenwei | +2/-2 | 1 个文件
cleanup_dist_env_and_memory() called torch._C._host_emptyCache() to free the host pinned-memory cache. That API is CUDA-specific and is not available on XPU, which triggered a warning on every shutdown. So this PR switches to the unified torch.accelerator.empty_host_cache() API. ## Test Result —
- 81bc196 #50910 — [Model Runner V2] Cache draft logits in model’s LM head dtype (#50910)
- 作者: Giancarlo Delfin | +170/-76 | 6 个文件
Summary Currently in MRV2 the draft logits are cached in FP32 so that they can be used for rejection sampling. The reason we cache in FP32 is because we are caching the temperature-applied logits, which need the FP32 numerical precision to prevent rounding errors. We can save 1/2 the memory by simply caching the draft logits in the data type of the model’s LM head (typically BF16), and then appl…
- f85c1d2 #51146 — K3: remove the add operation for megamoe path (#51146)
- 作者: Jee Jee Li | +27/-8 | 1 个文件
Test Result ### GSM8K - main branch - this PR —
- 9f31699 #50942 — [MoE] Align TRTLLM MXFP4 autotune buckets (#50942)
- 作者: Thien Tran | +6/-5 | 1 个文件
Use the shared FlashInfer MoE bucket helper so MXFP4 follows the same max-token and DP-aware convention as the other TRTLLM backends. Align with https://github.com/vllm-project/vllm/pull/47427, so that it will autotune up to 8192 tokens. Without this change, MXFP4 TRTLLM sets tune_max_num_tokens = moe_config.max_capture_size, which is only 512 -> bad for prefill. We can verify the behavior in vLLM…
- d36f24b #48069 — [KV Connector][Mooncake] Add tenant ID support to MooncakeStoreConnector (#48069)
- 作者: LZW | +149/-2 | 3 个文件
Add Mooncake tenant id support to MooncakeStoreConnector. Mooncake supports tenant-aware namespaces. This PR lets vLLM read an optional tenant_id from the Mooncake JSON config and forward it to MooncakeDistributedStore.setup(…, tenant_id=…) so different vLLM deployments can use separate Mooncake tenant namespaces. The default behavior remains backward compatible: - If tenant_id is absent, empt…
- b30291c #50390 — [EPD] Remove duplicate image preprocessing in EPD and enable preprocess on GPU (#50390)
- 作者: Tianyu Guo | +903/-44 | 22 个文件
Problem Two separate costs in an EPD deployment. The decode instance redoes work the encoder already did. The encoder instance runs the full HF image transform to produce the embedding; the decode instance then runs the identical transform over the same pixels. It does not need the result — the embedding already reaches it out of band through the EC connector, and the only thing it still ne…
- 38ebd97 #51083 — [ROCm] Relax MLA rope+cache test tolerances for bf16 (#51083)
- 作者: Rohan Potdar | +17/-5 | 1 个文件
tests/kernels/core/test_rotary_embedding_mla_cache_fused.py::test_concat_and_cache_mla_rope_fused fails on ROCm for bfloat16 against the hardcoded atol=0.001. On gfx942 this is 48 of 3153 parametrizations, all dtype=bfloat16 (fp16 and fp32 pass 100%). The failures are ~1 bfloat16 ULP, not a kernel bug. The fixture routes ROCm through the AITER Triton rope for fp16-consistent numerics; the fuse…
- 8779758 #51070 — [K3 Perf] Combine multiple all gather together for SP, 1.5~3x kernel level performance improvement (#51070)
- 作者: Wentao Ye | +21/-15 | 1 个文件
Originally, we do Now we only do a final allgather for optimization Acc covered in unit test Perf can be seen in this AI generated script And we can get
- 373fe8b #51078 — [MoE Refactor] Remove MoE legacy code (#51078)
- 作者: bnellnm | +3/-295 | 22 个文件
Remove deprecated MoE methods that are no longer needed due to oracle/modular kernel refactoring. Lint/CI ## Test Result — BEFORE SUBMITTING, PLEASE READ https://docs.vllm.ai/en/latest/contributing (anything written below this line will be removed by GitHub Actions)
- e76b71f #51179 — [CI Bug] Fix
pydantic_core._pydantic_core.ValidationError: Input should be a valid integer(#51179)- 作者: Wentao Ye | +4/-1 | 1 个文件
Introduced by https://github.com/vllm-project/vllm/pull/50940 Causing bug of https://buildkite.com/vllm/ci/builds/82455#019fd24f-5915-4781-9bc7-acb04aa9e04e This PR fixes the issue
- c2d8009 #51174 — [ROCm] Work around DeepEP teardown SIGSEGV in MoE test harness (#51174)
- 作者: Rohan Potdar | +12/-0 | 1 个文件
On ROCm, tests/kernels/moe/test_deepep_moe.py::test_deep_ep_moe fails with each spawned rank dying via SIGSEGV at process exit (no traceback). The MoE compute is correct; the crash is teardown-only — a ROCr use-after-free in HIP’s atexit handler, fixed upstream in https://github.com/ROCm/rocm-systems/pull/6942 but not yet in our base image’s ROCm build. Guard the DeepEP spawn worker so on ROCm it …
- 8f158d0 #44359 — [MoE] Share apply_moe_activation support metadata (#44359)
- 作者: Michael Goin | +425/-241 | 13 个文件
Centralize the activation capability and configuration contract for MoE backends that delegate to apply_moe_activation(). This adds apply_moe_activation_supported() and uses it from Triton, Marlin, Humming, and the applicable CUTLASS FP8/NVFP4/MXFP4/W4A8 experts instead of copying activation lists into each backend. It also adds the public, immutable ApplyMoEActivationConfig: - it is resolved once…
- 14e57ad #51067 — [Docker][KVConnector] Install mooncake from official wheels instead of a custom build (#51067)
- 作者: Zhewen Li | +26/-47 | 3 个文件
Release images install mooncake-transfer-engine from a custom wheel hosted on S3 (0.3.10.post2-0da9dfea3) rather than from PyPI. This PR removes that override and uses the official wheels, which greatly simplifies mooncake UX for anyone consuming vLLM outside our release images. The custom wheel was introduced in #42114 for two reasons, both of which have since been resolved upstream: 1. **WITH_NV…
- 613411a #51177 — [CI Bug] Fix
Chunked prefill is required for mamba cache mode 'align'.(#51177)- 作者: Wentao Ye | +1/-0 | 1 个文件
Fixes https://buildkite.com/vllm/ci/builds/82455#019fd24f-59b9-495b-a7c3-3cbac8155ee2
- 23c0f33 #51180 — [CI bug] Fix
Each KV cache group's real block_size must be divisible by has h_block_size(#51180)- 作者: Wentao Ye | +1/-4 | 1 个文件
Fixes https://buildkite.com/vllm/ci/builds/82455#019fd2a4-0c78-4fd7-8580-ac6d3ccaa8a8
- e8586b2 #51060 — [Build] Remove Ubuntu build-stage option from the CUDA dockerfile (#51060)
- 作者: Michael Goin | +42/-97 | 8 个文件
Implemented the CUDA manylinux-only build path. - Removed BUILD_OS and all Ubuntu build-stage branches from docker/Dockerfile. - Defaulted CUDA wheel builds to pytorch/manylinux2_28-builder:cuda13.0. - Updated x86_64/aarch64 CUDA 12.9 and 13.0 CI builders, documentation, generated versions metadata, and Docker dependency graph. - Preserved Ubuntu only for final runtime images. ## Test Result —
- 8543522 #51007 — [KV Offload] Support out-of-tree secondary tier managers via
module_path(#51007)- 作者: Ronen Schaffer | +153/-15 | 6 个文件
- Add config-driven out-of-tree loading to SecondaryTierFactory, matching the pattern already used by OffloadingSpecFactory (spec_module_path) and CachePolicyFactory (cache_policy_module_path) - Users can now specify custom SecondaryTierManager implementations purely via config without forking vLLM or running register_tier() before engine init - Document the feature in the usage guide ### Example …
- 62c5e21 #50507 — [KV Offloading] Support partial-tail prefix reuse with fine-grained prefix matching (#50507)
- 作者: Chauncey | +360/-44 | 2 个文件
see https://github.com/vllm-project/vllm/issues/45702 In hybrid Attention–Mamba models such as Qwen3.6, physical KV blocks are often very large to accommodate the Mamba state. Even with a smaller prefix_match_unit, native offloading previously could only store and restore complete physical blocks. As a result, reusable prefix tokens already computed near the end of a block were lost. This change e…
- 397094d #49990 — Resolve revision to commit_hash once per model load, via huggingface_hub’s
resolve_revision(#49990)- 作者: Lucain | +84/-6 | 9 个文件
TL;DR: resolve revision and remote code_revision commit hashes only once via huggingface_hub’s new resolve_revision, preventing repeated downstream resolution and avoiding concurrency issues if a repository is updated while a model is loading Disclaimer: AI-assisted PR, heavily human-reviewed/amended. I’m maintainer of the huggingface_hub which handles all the download/cache system. cc @hme…
- 5bf3260 #50448 — [Rust Frontend] Deduplicate request preprocessing for
/tokenize(#50448)- 作者: Sage | +170/-50 | 5 个文件
/tokenize currently repeats prompt handling already used by inference. This can make its token IDs drift from what generation receives, especially after chat rendering or multimodal prompt expansion. This pr expose and reuse the chat and text request processors from vllm-project/vllm#49045.
- fd859d2 #44956 — [KV Connector][Mooncake] Add store group semantics (#44956)
- 作者: Schatten | +531/-3 | 2 个文件
Mooncake added optional grouped object semantics in kvcache-ai/Mooncake#2180, motivated by the group-semantics discussion in kvcache-ai/Mooncake#2127. vLLM’s MooncakeStoreConnector currently writes each physical Mooncake object independently, even when multiple physical objects belong to the same logical vLLM KV cache entry. This PR adds an opt-in vLLM-side integration so related Mooncake objects …
- 999dd8b #50526 — [XPU] Alias is_current_stream_capturing to XPU in cuda wrapper (#50526)
- 作者: Sundaresan G | +1/-0 | 1 个文件
The XPU backend aliases torch.cuda.* APIs to their torch.xpu.* equivalents in _torch_cuda_wrapper so that CUDA-graph-capturing code paths work on XPU. The graph-related aliases (torch.cuda.graph, CUDAGraph, graph_pool_handle) were already set, but torch.cuda.is_current_stream_capturing was missing, so any code that guards on stream capture hit an AttributeError on XPU. This adds the missing alias …
🐛 Bug Fix
- 46e6a83 #51227 — [Bugfix][KV Offload] Clean up resources after initialization failure (#51227)
- 作者: AlexHuang | +85/-53 | 2 个文件
[Bugfix][KV Offload] Clean up resources after initialization failure Prevent /dev/shm/vllm_offload_*.mmap and tier resources from leaking when KV offloading initialization fails. - Clean up the scheduler mmap when primary-tier construction fails. - Shutdown already-created secondary tiers in reverse order when a later tier fails, then continue with primary cleanup. - Clean up CPU/tiering worker …
- 5fba75a #51249 — [Bugfix][Model] Add missing fused_qkv_a_proj to Kimi-Linear packed_modules_mapping (#51249)
- 作者: yjz | +1/-0 | 1 个文件
While testing a Kimi-K3 W4A8 quantized checkpoint, we noticed that several Linear layers in the checkpoint are intentionally left unquantized (bf16) rather than packed as W4A8 — attention projections among them. Digging into why one of those bf16 layers was being treated as quantized led to this bug. KimiLinearModel.packed_modules_mapping (added in #50500) lists the fused params for MoE/dense MLP …
- ef2615c #51116 — [Bugfix][KV Offload] Fall back when MADV_POPULATE_WRITE is unsupported (#51116)
- 作者: AlexHuang | +197/-6 | 2 个文件
Fixes #44888 ## Why this PR SharedOffloadRegion.init unconditionally issues madvise(MADV_POPULATE_WRITE, …) on its shared mmap. On Linux < 5.14 (RHEL 8 ships 4.18) the syscall returns EINVAL, so any vLLM startup that touches TieringOffloadingSpec aborts before KV offloading can register. This PR probes the advice once after mmap creation and falls back to a no-mutation NumPy read-modify-write …
- 872fd59 #51002 — [Bugfix][LoRA] Guard TrtLlm BF16 MoE LoRA gate on activation type (#51002)
- 作者: anhtra3889 | +10/-12 | 1 个文件
Fix BF16 MoE + LoRA backend selection for non-gated models (e.g. NemotronH). The LoRA path in select_unquantized_moe_backend() returns early via _trtllm_bf16_lora_supported() before the modular-kernel oracle runs, so activation / act_and_mul checks on TrtLlmBf16LoRAExperts were skipped. Non-gated experts were routed into the gated-only FlashInfer kernel and crashed at startup during profile_run wi…
- 5370461 #51108 — [BugFix][KV Cache] Fix hybrid prefix caching with hidden-state extraction (#51108)
- 作者: Canlin Guo | +82/-1 | 2 个文件
Reproduce the error below in the main branch. Fix https://buildkite.com/vllm/ci/builds/82382/canvas?sid=019fce99-6506-4c56-a39c-5bde8f023b8d&tab=output error introduced by #50991. After this PR: ## Test Result —
- 2e35c52 #51038 — [Bugfix][Quantization] Fix MXFP4 conversion for FlashInfer CUTLASS (#51038)
- 作者: Gabriel Wu | +137/-2 | 2 个文件
The MXFP4 backend selector can choose FlashInfer CUTLASS for standard checkpoints, but convert_weight_to_mxfp4_moe_kernel_format has no corresponding branch and raises ValueError after model loading. This change: - reorders standard fused gate/up weights, scales, and bias from [w1; w3] to the [w3; w1] layout expected by FlashInfer CUTLASS; - applies the backend-specific MXFP8 scale interleave or B…
- c0202c5 #48341 — [Bugfix][Spec Decode] Auto-enable async scheduling for draft models (#48341)
- 作者: JooHo Lee | +25/-39 | 3 个文件
- Enable asynchronous scheduling by default for draft-model speculative decoding. - Add configuration regression coverage for the default behavior. - Remove the issue-specific xfails and exception now that the affected runtime paths pass normally. VllmConfig already permits asynchronous scheduling when it is explicitly enabled with draft_model, but the default-resolution path classified draft_mode…
- 8217171 #50890 — [BugFix][Pooling] Skip weight-prefix probe when model has WeightsMapper (#50890)
- 作者: Li Juqi | +108/-2 | 2 个文件
[BugFix][Pooling] Skip weight-prefix probe when the model owns a WeightsMapper ModelForPooling.load_weights always runs a generic target_prefix probe (candidate_prefixes = ["", “model.”]) before delegating to the wrapped model. To stay correct under buffer-reusing iterators (see #39650), every scanned tensor is .clone()’d into seen_weights. That probe never matches checkpoints whose keys only li…
- f84df12 #48413 — [Bugfix][MM] Fix MiniCPM-V placeholder replacement and image processor loading on Transformers v5 (#48413)
- 作者: Yunzhu Lu | +218/-10 | 4 个文件
This PR fixes two independent but related issues that block MiniCPM-V models (2.5 / 2.6 / 4.0 / 4.5) from working with Transformers v5 in vLLM: 1. Placeholder mismatch: Multi-modal placeholder replacement fails at runtime with found 0 prompt placeholders. 2. Wrong Image processor reuse: When multiple MiniCPM-V checkpoints are loaded in the same process (e.g. serial pytest runs), later models incor…
- b50fdeb #39935 — [Bugfix] Fix level-2 sleep/wake/reload with enable_lora=True (#39935)
- 作者: Silen Naihin | +252/-4 | 11 个文件
Fixes #39934. Fix level-2 sleep/wake/reload for LoRA-enabled models. Without this fix, reload_weights() crashes because LoRA wrapping moves parameters under base_layer (checkpoint names no longer match), and LoRA GPU state (stacked tensors, adapter registry, TP>1 logits mapping) is left undefined after level-2 sleep discards its memory. Three changes, keeping standard PyTorch module-tree semantics…
- 8cfa01c #49212 — [BugFix] Dense multinode DP rescope with regression test (#49212)
- 作者: aoshen02 | +38/-7 | 3 个文件
This recreates #48626, which GitHub automatically closed after the source fork was accidentally converted to a private standalone repository. The original PR could not be reopened after that repository left the upstream fork network. This PR supersedes #39405 and carries forward Andy Lo’s original fix for port conflicts in dense models with multi-node data parallelism. The original two commits rem…
- 47a4e41 #50183 — [Bugfix][Spec Decode] Fix NaN handling in rejection sampler tl.argmax (#50183)
- 作者: gabriel-peracio | +51/-0 | 2 个文件
tl.argmax on an all-NaN block returns an out-of-range block index (pointing into the padded region beyond num_blocks - 1). That index is used to load from the local-argmax tensor, producing an out-of-bounds read and an illegal memory access downstream in the rejection sampler. Two kernels in vllm/v1/worker/gpu/spec_decode/rejection_sampler_utils.py are affected: - _compute_global_target_argmax — t…
- 7c77868 #51125 — [Bugfix] Size and iterate w13 by shard count for non-gated MoE (#51125)
- 作者: aoshen02 | +104/-80 | 16 个文件
Non-gated MoE (is_act_and_mul=False, e.g. NemotronH’s relu2_no_mul) fuses only the up projection into w13, so w13 holds a single intermediate_size_per_partition shard rather than two gate/up shards. 13 MoE quantization methods already conditionalize on this — modelopt.py, compressed_tensors_moe_*, flashinfer, unquantized_fused_moe_method.py, … — each with its own local w13_num_shards = 2 if se…
- 482c613 #51093 — [Bugfix][Humming] Preserve ModelOpt FP8 weight dimensions (#51093)
- 作者: Netanel Haber | +25/-0 | 2 个文件
ModelOpt FP8 transposes serialized weights from [N, K] to [K, N] before dispatching to the selected linear kernel, but the replacement Parameter does not preserve the weight’s dimension metadata. Humming then falls back to interpreting the tensor as [N, K]. GSM8K with and w…
- 9833aa5 #50358 — [Bugfix] Fail fast with a clear error when CPU offload region exceeds available space (#50358)
- 作者: AlexHuang | +95/-6 | 2 个文件
[Bugfix] Fail fast when CPU KV offload exceeds /dev/shm Fixes #46949. Oversized CPU KV offload regions currently fail in mmap.madvise(MADV_POPULATE_WRITE) with an opaque OSError: [Errno 14] Bad address. This PR checks /dev/shm capacity first and reports the requested and available sizes with actionable configuration guidance. After #50094, the default CPUOffloadingSpec also uses SharedOffloadReg…
- 08b8613 #48929 — [Bugfix][Model] Fix MiniMax-M3 NVFP4 inference correctness (#48929)
- 作者: Gabriel Wu | +100/-31 | 2 个文件
Fix two MiniMax-M3 correctness issues exposed by the NVFP4 checkpoint. First, the routed experts use packed SWIGLUOAI_UNINTERLEAVE with model-specific alpha, beta, and clamp values. FlashInfer CUTLASS already supports this math, but the vLLM adapter neither advertised the packed activation nor forwarded all three parameters. Marlin similarly replaced missing quant-config alpha/beta values with pla…
- cd930c8 #38771 — [Bugfix] Fix MLA kv_b_proj activation dtype with Marlin FP8 (#38771)
- 作者: Jacob Zhang | +24/-25 | 1 个文件
Fixes #38658. This PR fixes an MLA prefill dtype bug when FP8 weights are served through the Marlin path on GPUs without native FP8 support (sm < 89). On affected GPUs, Marlin repacks FP8 weights into torch.int32. In vllm/model_executor/layers/attention/mla_attention.py, compute_prefill_context() was using self.kv_b_proj.weight.dtype to determine how to cast kv_c_normed before passing it to kv_b…
- a9b39d6 #51153 — [Bugfix] Enable chunked prefill for qwen3.5-0.8B ppl test (#51153)
- 作者: music-dino | +3/-0 | 1 个文件
PR #50991 enabled prefix caching by default and changed the default mode to align, and updated several hybrid model tests to enable chunked prefill to go along with the new defaults. It missed models/language/generation_ppl_test/test_qwen.py::test_ppl[model_info2] so the test failed in the Language Models Test (PPL) on the nightly. This PR enables chunked prefill for the failing hybrid model test …
- beca88e #51131 — [BugFix][K3] Skip moe_intermediate padding when EP is enabled (#51131)
- 作者: Ziming Huang | +1/-1 | 1 个文件
FIX https://github.com/vllm-project/vllm/issues/51124 ## Test Result —
- f5cd862 #50649 — [ROCm][Bugfix] Kimi-K3 Fix KDA NaN on mixed batches and racy autotune config (#50649)
- 作者: kliuae | +556/-7 | 3 个文件
For Kimi-K3, there are two correctness bugs on the ROCm Kimi-K3 KDA path. - KimiGatedDeltaNetAttention._forward sends the decode sequences of a mixed decode+prefill non-speculative batch through the KDA chunk kernel chunk_kda_with_fused_gate. The chunk kernel treats each decode as length=1 sequence, and returns NaN for them. This PR separates the decode out of the mixed batch and pass them to fuse…
🔩 Misc
- e07532b #50981 — [MISC][Bench] refactor throughput and reuse serve’s get samples (#50981)
- 作者: Jared Wen | +228/-231 | 3 个文件
#50849 #50838 vllm bench serve’s datasets is the superset of vllm bench throughput let throughput reuse the get_samples like serve, we can make these two share the same datasets without manually adding each dataset to throughput or serve and we can eliminate the tons of if else in throughput for maintenance ## Test Result —
- bc37fc9 #50879 — Revert [Misc] Avoid importing
nixl_epon everyvllm serveconfig (#50879) (#51176)- 作者: fxmarty-amd | +23/-24 | 1 个文件
Avoid importing nixl_ep all the time we boot up vllm. Instead, only do it lazily when needed (dpep configuration). Right now if you do vllm serve … import_utils will try to resolve nixl_ep, and in this case even log “unrelated” import failures as this configuration can then go on and run just fine, as no DPEP is needed (I have no libcudart 12 on this machine, but it only matters if it’s actual…
⚡ Performance
- 2dfb8ba #50230 — [Perf][CUDA] Programmatic dependent launch for the DSA decode kernels (#50230)
- 作者: xiaozhoupy | +54/-10 | 2 个文件
What this does Chain the DSA decode kernels with programmatic dependent launch so back-to-back small kernels overlap their launch latency: - the fused norm/rope and fused-q Triton kernels get a USE_PDL constexpr with gdc_wait() / gdc_launch_dependents(), launched with launch_pdl - the NVFP4 quant kernels switch to cudaLaunchKernelEx with programmatic stream serialization, guarded on __CUDA_ARCH…
- b92352c #50992 — [Perf][KV Offload] Avoid quadratic ARC batch eviction (#50992)
- 作者: MINJUN GIL | +132/-21 | 2 个文件
ARC batch eviction repeatedly scans its internal cache lists from the beginning for each block selected. As a result, evicting many blocks can require quadratic work. ## Fix Keep monotonic iterators over those lists while collecting candidates, so each entry is visited at most once. Cache mutations remain deferred until all requested candidates are found, preserving atomic eviction. This also pres…
🖥️ Kernel
- 8060260 #49932 — [Linear] [Kernel] add block-wise scaled_mm (#49932)
- 作者: zofia | +171/-0 | 3 个文件
This adds a native, dependency-free torch._scaled_mm block-wise FP8 backend. It needs no extra libraries or per-shape tuning (unlike CUTLASS/DeepGEMM/Triton), and stays competitive-to-best in the compute-bound regime (M≥1024, up to 1.09x faster than the next-best backend on H20). Add a block-wise FP8 _scaled_mm linear kernel (BlockWiseTorchFP8ScaledMMLinearKernel) that routes DeepSeek-style block …
- 66b3c0e #50294 — [Kernel][Model] Optimize FA4 mm_prefix range lookup (#50294)
- 作者: Li Juqi | +819/-45 | 5 个文件
[Kernel][Model] Gemma4: optimize FA4 mm_prefix range lookup and CuTe JIT stability Gemma4 multimodal models can use FA4 as the backend for all layers, so that both sliding-attention layers (head_dim=256) and global/full layers (global_head_dim=512) use one attention backend. This is required for the vision mm_prefix / PrefixLM bidirectional mask path, and avoids mixing FlashAttention and Triton …
🧪 CI/Tests
- 2a0323b #51160 — [XPU][Test] Support MultiConnector accuracy testing on XPU (#51160)
- 作者: liuzhenwei | +20/-5 | 2 个文件
Support MultiConnector accuracy testing on XPU KV_BUFFER_DEVICE=xpu NUM_CONCURRENT=1 bash tests/v1/kv_connector/nixl_integration/run_multi_connector_accuracy_test.sh ## Test Result —
- 71f975a #50480 — [ROCm][CI] Add MLA decode accuracy and determinism tests (#50480)
- 作者: Aarushi Jain | +261/-0 | 2 个文件
Test the production BF16 decode path through rocm_aiter_ops.mla_decode_fwd which is the hot path for DeepSeek-V3/V4 inference on MI300/MI355. Covers: - Smoke test (shape, dtype, finite, non-zero) - Accuracy vs PyTorch reference (absorbed MLA formulation) - Parametrized sweep: nhead in {16, 128}, batch in {1, 4, 16}, seq in {16, 256} - Bitwise determinism across 4 runs Uses atol=0.01, rtol=0.0, pas…
- 4282fe5 #49375 — [ROCm][CI] Add More AITER quantization/MoE kernel tests (#49375)
- 作者: Micah Williamson | +2335/-2 | 5 个文件
This PR expands ROCm kernel test coverage for AITER quantization and fused-MoE paths. Previously these ROCm-specific code paths (FP8, MXFP4/FP4, and AITER fused MoE) had little to no targeted kernel-level testing on AMD hardware.
- 65addac #51173 — [ROCm][CI] Keep rocprofiler-sdk out of DeepEP HT MoE test workers (#51173)
- 作者: stefankoncarevic | +7/-0 | 1 个文件
On ROCm, every high-throughput variant of test_deep_ep_moe fails with SIGSEGV (28 of 28), while every low-latency variant passes. The crash is not in the test and not in DeepEP: the test body completes, and the process then dies during teardown inside the ROCm HSA runtime, where AqlQueue’s destructor writes to signals whose backing memory has already been unmapped. That destructor is only reached …
- e6d67fd #51074 — [CI] Prune PyTorch Fullgraph Test (#51074)
- 作者: Michael Goin | +25/-124 | 3 个文件
Test Result —
- a3b8675 #51095 — [CI] Fix CI authorization notification fallback (#51095)
- 作者: Kevin H. Luu | +97/-13 | 5 个文件
- resolve approval notifications from the source workflow’s head commit when workflow_run.pull_requests is empty or unrelated - let a ready label recover a missed approval notification while retaining marker-based deduplication - give approval-triggered notification runs a unique concurrency fallback when no PR number is attached - run add_label_automerge, notify-ci-authorized, and record-ci-appro…
- 96e333e #51127 — [CI] Run control-plane workflows on vLLM runners (#51127)
- 作者: Kevin H. Luu | +2/-2 | 2 个文件
- run the pre-commit pre-run-check gate on the vllm-runners self-hosted runner group - run the PR-comment CI broker on the same vllm-runners group - keep the existing self-hosted pre-commit runner labels unchanged Both jobs now use the repository’s established selector: ## Why The organization was at its 20/20 standard GitHub-hosted concurrency limit. For the pre-commit workflow, waiting on the Gi…
✨ New Feature
- 2217035 #51089 — [Feature] Parse request priority from HTTP header (#51089)
- 作者: Chauncey | +81/-8 | 8 个文件
FIX https://github.com/vllm-project/vllm/issues/51023 Parse request priority from HTTP header ## Test Result —
🔧 Refactor
- 811622c #50066 — [Refactor][PCP] Make PCPManager construction extensible (#50066)
- 作者: Qiu | +9/-2 | 2 个文件
maybe_build_pcp_manager currently hard-codes both PCPManager.validate_config and PCPManager(…). This prevents callers from using a specialized PCPManager subclass through the existing construction path, forcing them to duplicate the factory logic or replace global symbols. This change adds an optional manager class argument, defaulting to PCPManager, and uses that class for both validation and c…
🦀 Rust Frontend
- d4da0c5 #51045 — [Model][Frontend] Add Ling 3.0 Flash BF16, MTP, and parser support (#51045)
- 作者: zexplorerhj | +2070/-1 | 15 个文件
Add BF16 inference support for inclusionAI/Ling-3.0-flash, including MTP speculative decoding and Ling3 reasoning/tool-call parsing. ## Implementation - Add the Bailing V3 MLA/KDA hybrid MoE model implementation and its MTP draft model. - Register the base and MTP architectures, speculative-model config rewrite, architecture config converter, and model registry test entries. - Support the model’s …