共 70 个 commit,涉及 364 个文件,+14810/-1243 行变动。
概要
| 统计项 | 数值 |
|---|---|
| Commit 数 | 70 |
| 变更文件 | 364 |
| 新增行数 | +14810 |
| 删除行数 | -1243 |
Commit 列表
🐛 Bug Fix
- 96acd47 #52122 — [Bugfix][MiniCPM-V] Fix AssertionError in get_dummy_mm_data when passing VideoDummyOptions to _get_dummy_images (#52122)
- 作者: Qiming Zhang | +18/-2 | 1 个文件
Issue: pytest tests/lora/test_minicpmv_tp.py::test_minicpmv_lora raises AssertionError on non-CUDA platforms (e.g., XPU). Root Cause: Commit 9a276d6375 added a runtime assertion to _get_dummy_images in dummy_inputs.py: assert overrides is None or isinstance(overrides, ImageDummyOptions) However, MiniCPMVDummyInputsBuilder.get_dummy_mm_data in minicpmv.py had always been passing video_overr…
- 64ca614 #47692 — [Bugfix] Fix
--data-parallel-start-rank 0being treated as unset increate_engine_config(#47692)- 作者: Ali Jaseem | +21/-2 | 2 个文件
EngineArgs.create_engine_config uses Python truthiness (if self.data_parallel_start_rank) instead of is not None to detect whether –data-parallel-start-rank was explicitly set. Since 0 is a valid, meaningful starting rank (the node owning the first slice of global DP ranks), an explicit –data-parallel-start-rank 0 is silently treated identically to “not specified.” This causes data_parallel_hybr…
- 2d24355 #52030 — [Bugfix] Fix packed GDN decode launch for large batch-head grids (#52030)
- 作者: Michael Goin | +33/-3 | 2 个文件
Avoid a CUDA launch failure in packed GDN decode when batch_size * num_value_heads exceeds the maximum CUDA grid Y/Z dimension of 65,535. The existing launch is preserved for normal sizes. Only overflowing cases use a split (value_tiles, value_heads, batch) grid. ## Test Result - Verified the failing Qwen shape (B=1024, HV=64, K=V=128) launches successfully. - Running vllm serve mgoin/Qwen3.8-2.4T…
- 170592a #52172 — [Bugfix] Disable sequence parallelism for Dots3 NOTE (#52172)
- 作者: 范裕达 | +3/-5 | 1 个文件
Why The DeepSeek V3.2 sequence-parallel refactor changed the inherited forward paths to use use_sequence_parallel. Dots3 NOTE uses custom model and decoder initializers and has not adopted the new sequence-parallel execution path. This causes serving to fail during KV cache profiling with: AttributeError: ‘Dots3NoteModel’ object has no attribute ‘use_sequence_parallel’ ## What changed Explicitl…
- d0ae25e #52021 — [Bugfix] Preserve Anthropic disable_parallel_tool_use (#52021)
- 作者: Taneem Ibrahim | +4/-0 | 2 个文件
Anthropic’s disable_parallel_tool_use was silently discarded, leaving the converted OpenAI request with parallel_tool_calls=True. This preserves the field and maps it to the existing inverse OpenAI setting. ## Reproducer On Main On this branch ## Test Plan and Results
- 399f974 #50874 — [Bugfix][R3] Size monolithic routing replay buffer for DP (#50874)
- 作者: TomerBN-Nvidia | +30/-5 | 2 个文件
Fix routing-replay capture for the FlashInfer monolithic MoE kernel under naive data parallelism, including padded sequence-parallel shards when expert parallelism is enabled. Two related assumptions fail in a TP2/DP2 deployment: 1. Replay buffer capacity. max_num_tokens is a per-rank scheduler limit, while the naive dispatch path all-gathers rank-local batches before invoking the monolithic k…
- 2ac1f68 #51256 — [BugFix] Reserve the bonus query slot in DFlash scheduling budget (#51256)
- 作者: kx | +59/-10 | 2 个文件
The existing generic parallel-drafting calculation only reserves K - 1 slots for DFlash, leaving the scheduling budget short by one slot per request. For example, with: max_num_batched_tokens = 2048 max_num_seqs = 256 num_speculative_tokens = 8 the previous calculation allowed: max_num_scheduled_tokens = 2048 - 7 * 256 = 256 However, a full DFlash batch may require: 256 * (8 + 1) = 2304 query toke…
- 79f3183 #52058 — [Bugfix] Bound KV block zeroing launch geometry (#52058)
- 作者: Lucas Wilkinson | +86/-37 | 2 个文件
Fix the KVBlockZeroer launch overflow reproduced on main nightly #83443, at commit 3e372c5ff2: DeepSeek-V4 combines 181 KV segments with 9,344- and 292-element pages. The old zeroer selected the largest common power-of-two divisor, so the 292-element page forced every segment to use four-element chunks. Zeroing 6,870 blocks therefore flattened to: That overflows the signed launch dimension passed …
- 15227b9 #51614 — [Bugfix][KV Offload] Emit self-describing CPU events at KV-group block granularity (#51614)
- 作者: Ziqi Fan | +59/-16 | 2 个文件
Fix self-describing CPU KV events for hybrid KV-cache layouts where request hashes are computed more frequently than the full-attention group’s block size. For example, DeepSeek V4 may use: - tokens_per_hash = 4, derived from the GCD of its KV-group block sizes - tokens_per_block = 256 for the full MLA group The offloading event tracker currently publishes every raw 4-token hash and sets BlockStor…
- b369f10 #51821 — [Bugfix][ROCm][CI] Restore the DeepSeek-V4 input GEMM override point (#51821)
- 作者: stefankoncarevic | +8/-1 | 1 个文件
The GSM8K accuracy job for amd/DeepSeek-V4-Flash-NVFP4 on gfx950 reports 0.0000 against a threshold of 0.92. It is not a near miss: the server comes up clean and then answers every one of the questions with an unparsable run of repeated tokens, so the failure is in the numerics rather than in the harness. A bisect over the window in which the job turned red lands on 79c865b838e34f7a98a936771284773…
- 5936bac #51218 — [Bugfix] Report FULL_ATTENTION for uniform-base UniformTypeKVCacheSpecs groups instead of UNKNOWN (#51218)
- 作者: Yifan Jiang | +121/-0 | 2 个文件
Problem get_kv_cache_spec_kind() returns KVCacheSpecKind.UNKNOWN for a UniformTypeKVCacheSpecs whose members have more than one inner kind: But such a group is only ever formed when all members share one registered uniform_type_base_spec — that is the merge condition itself (UniformTypeKVCacheSpecs.is_uniform_type → KVCacheSpec.is_uniform_with_collection). So UNKNOWN discards information …
- 98f86b9 #50017 — [ROCm] [bugfix] Chunked prefill paged decode masked load perf (#50017)
- 作者: afriedri | +34/-20 | 1 个文件
Observed 10-15% latency performance regression for serving Qwen/Qwen3-30B-A3B-Thinking-2507; trace revealed that kernel_paged_attention_2d was the culprit, taking 1.33x as long on average on v0.25.0 vs v0.24.0. #47305 fixed correctness but introduced performance drop because of universally applied masking. Edited masking so that it is only enacted during last block, when token id can be greater th…
- 34735ac #52028 — [Bugfix] Pin DeepEP by its full commit hash (#52028)
- 作者: Tyler Michael Smith | +3/-1 | 1 个文件
tools/ep_kernels/install_python_libraries.sh pins DeepEP by a 10-character abbreviation: An abbreviated object ID is not a ref, so it cannot be fetched directly. GitHub serves any complete commit via allowAnySHA1InWant, but an abbreviation is never a valid want: Cloning the whole repository and checking out afterwards resolves the abbreviation locally, which is why this goes unnoticed in the com…
- e62abc3 #49139 — [Bugfix][Kernel] Fix persistent top-k histogram reuse after short rows (#49139)
- 作者: xiao feng | +47/-3 | 2 个文件
Fix a correctness bug in persistent_topk when one persistent CTA group processes a radix row (seq_len > 32768), followed by a short or medium row (seq_len <= 32768), and then another radix row. The kernel previously used the outer row iteration counter to rotate its triple-buffered radix histograms. Short rows advanced that counter without executing radix_topk, causing the next radix row to reuse …
- 19d61b1 #51359 — [Bugfix] Initialize DeepGemmQuantScaleFMT oracle lazily; bound QuantFP8 UE8M0 packed path to group_size 128 (#51359)
- 作者: kiroxu | +30/-2 | 3 个文件
Fix a crash and a latent kernel-contract violation in QuantFP8’s DeepGEMM UE8M0 fast path, found by running the FP8 kernel suite on an RTX PRO 6000 (SM120). DeepGemmQuantScaleFMT.from_oracle() asserts its cache is populated, but the cache is only filled as a side effect of _lazy_init() running for some other DeepGEMM wrapper (it’s called at the end of _lazy_init, #30898). A QuantFP8 built with an …
- d20a031 #51980 — [Bugfix][ROCm][MoE] Update AITER MXFP4 W4A16 tests to the renamed expert_mask (#51980)
- 作者: stefankoncarevic | +2/-2 | 1 个文件
Fixes two errors in tests/kernels/moe/test_rocm_aiter_moe.py on gfx950: #49758 renamed that keyword from expert_map to expert_mask and updated the production callers (AiterExperts.apply, quark), but not these two test call sites. This renames them; both pass None, so nothing else changes. The other expert_mask= uses in the same file call torch.ops.vllm.rocm_aiter_fused_moe, which has always used t…
- 6ea6c42 #51997 — [Bugfix] Bound Anthropic stop sequences (#51997)
- 作者: Taneem Ibrahim | +6/-2 | 1 个文件
PR #51447 bounded OpenAI stop parameters with VLLM_MAX_STOP_STRINGS. Anthropic requests remained unbounded, despite being converted internally into a bounded ChatCompletionRequest. Consequently, an Anthropic request with five stop sequences passed ingress validation but failed during conversion, returning HTTP 500. This change applies the same limit at the Anthropic API boundary. No duplicate open…
- 8e95890 #51843 — [Bugfix] Disable fine-grained prefix-cache hits for incompatible hybrid KV layouts (#51843)
- 作者: Michael Goin | +96/-3 | 3 个文件
Fine-grained prefix-cache hits are enabled for hybrid models containing Mamba align groups. However other groups, such as a sliding-window DSpark drafter, may use KV cache managers that only support block-aligned lookups. This previously caused an assertion when prefix caching was enabled. This change: - Enables fine-grained hits only when every KV cache manager supports the required lookup granul…
- 4c51ceb #51139 — [Bugfix][Multimodal] Invalidate retained PyNvVideoCodec decoder after failure (#51139)
- 作者: dmai-afk | +134/-1 | 3 个文件
Title [Bugfix][Multimodal] Invalidate retained PyNvVideoCodec decoder after failure # Body Fixes #51138 - invalidate a retained PyNvVideoCodec decoder slot when a borrowed operation exits with an error - clear stale slot state before constructing a replacement decoder - publish the replacement decoder only after construction succeeds - add CPU-only regression coverage for borrow-time invalidatio…
- 86e2ab5 #51120 — [Bugfix][Frontend] Return 400 for invalid PyNvVideoCodec video input (#51120)
- 作者: dmai-afk | +110/-5 | 3 个文件
- translate PyNvVideoCodec.PyNvVCException raised while opening malformed video into a sanitized ValueError - rely on vLLM’s existing ValueError handler to return BadRequestError/HTTP 400 instead of HTTP 500 - preserve the native exception as the cause and leave unrelated decoder exceptions unchanged ## Root cause A malformed user-supplied video reaches PyNvVideoCodec.SimpleDecoder during multimod…
- e60f3c4 #46845 — [Bugfix] Fix MiniMax-M3 compressed-tensors FP8 MoE SwiGLU params (#46845)
- 作者: Tan Pin Siang | +34/-0 | 2 个文件
This PR fixes MiniMax-M3 accuracy when served from a compressed-tensors W8A8 FP8 MoE checkpoint. MiniMax-M3 uses hidden_act=“swigluoai” with non-default SwiGLU parameters: - swiglu_alpha=1.702 - swiglu_beta=1.0 - swiglu_limit=7.0 The compressed-tensors FP8 MoE path already forwarded swiglu_limit, but did not forward swiglu_alpha or swiglu_beta into make_fp8_moe_quant_config. As a result, MiniMax-M…
- aeece10 #51928 — [Bugfix][XPU] Run GDN attention as eager break under breakable cudagraph (#51928)
- 作者: Huanxing | +2/-1 | 1 个文件
When breakable graph, Qwen3.5 on XPU produced garbage output. This PR is fix the GDN (Gated Delta Net) recurrent attention on XPU. The XPU GDN attention code reads and writes the conv and ssm state caches in place using data-dependent metadata. Under breakable-cudagraph capture the capture-time state addressing is frozen into the graph, so replay corrupts the recurrent state and hybrid models (e.g…
📦 Other
- 80d6d55 #52147 — Standardise weight tying on
ParallelLMHead.tie_weights(#52147)- 作者: Harry Mellor | +191/-214 | 56 个文件
vLLM expresses tied word embeddings in three different ways. Only one of them, self.lm_head = self.lm_head.tie_weights(embed_tokens), dispatches through quant_method.tie_weights. This PR converts the other two so tying is expressed one way everywhere. - self.lm_head.weight = embed_tokens.weight (33 sites) bypassed the quant method entirely, so it was wrong for quant methods that repack. The Parall…
- f3c1638 #51653 — [ROCm] Enable V2 model runner for Kimi-K3 on ROCm (#51653)
- 作者: vllmellm | +0/-16 | 1 个文件
Kimi-K3 on ROCm was gated from using V2 model runner. After validation using the up-to-date upstream. V2 model runner is working as expected. Command to start Kimi-K3 on mi355x ## Test Result Tested with lm-eval: | runner | run 1 | run 2 | |——–|——-|——-| | V1 | 96.66% | 97.04% | | V2 | 97.57% | 96.89% | —
- 95c9144 #42662 — [LoRA][Gemma4] Support vision tower LoRA (#42662)
- 作者: linitra24 | +309/-48 | 10 个文件
This PR adds the remaining LoRA plumbing needed for Gemma4 multimodal LoRA support. After #43798, Gemma4-MM vision linear layers are already converted through the Transformers backend path, so this PR no longer reimplements the Gemma4 vision tower. Instead, it focuses on the runtime LoRA mapping and token-counting pieces needed by Gemma4 image/video/audio inputs. Main changes: - Add a multimodal L…
- 8e1131e #52114 — [Model] [Quantization] Add Ling hybrid MXFP4 routed experts support (#52114)
- 作者: zexplorerhj | +48/-4 | 2 个文件
Add support for Ling checkpoints that use hybrid quantization: block FP8 for dense and shared-expert projections, and MXFP4 for routed experts. This change reads Ling-specific quantization metadata and remaps routed-expert scale names to the convention expected by Mxfp4MoEMethod for both the main and MTP models
- c4e9692 #50534 — [XPU] Add tuned Mamba SSU configs for Intel Arc Pro B70 (#50534)
- 作者: pmanczak | +247/-12 | 5 个文件
Add tuned selective_state_update configs for Intel Arc Pro B70 Graphics, and fix benchmarks/kernels/benchmark_selective_state_update.py so it runs on an XPU-only build. The config directory holds AMD and NVIDIA devices only, so Mamba/hybrid models on Intel GPUs fall back to the heuristic. Four files, two shapes: | shape | cache_dtype | exercised by | | — | — | — | | headdim=128,dstate=25…
- 5fee0a8 #51998 — chore: Upstream Cohere parser fixes + tests (#51998)
- 作者: JasonCohere | +533/-54 | 10 个文件
Adding some local fixes for Cohere parsers alongside corresponding tests
- 37c3bdf #50221 — fix(security): enforce audio decode duration limit in NanoNemotronVL (#50221)
- 作者: Juan Pérez de Algaba | +63/-2 | 2 个文件
The _extract_audio_from_videos method called load_audio_pyav without max_duration_s, allowing a small compressed video to decompress into gigabytes of PCM and crash the server via OOM. Pass VLLM_MAX_AUDIO_DECODE_DURATION_S to match the safeguard already used by AudioMediaIO. This should be merged only when https://github.com/vllm-project/vllm/pull/49948 branch) is merged
- b8baa31 #49458 — Hardware-agnostic model definition via HF transformer backend (1/N) (#49458)
- 作者: Thomas Ortner | +691/-5 | 10 个文件
This PR is an alternative approach to realize hardware-agnostic model definitions based on the HF transformer backend. In particular, the idea is that the modeling code of tail models resides in HF transformers and can be executed in vLLM through the help of the transformers backend, i.e., –model-impl transformers. The way this backend currently works is by replacing / patching particular layers …
- 903da60 #52134 — [Docs] Fix
WhisperEncoderLayer.forwarddocstring indots3_note(#52134)- 作者: Harry Mellor | +18/-9 | 1 个文件
The docs build emits griffe warnings for vllm/models/dots3_note/nvidia/audio_encoder.py: The WhisperEncoderLayer.forward docstring was inherited from the upstream HF Whisper implementation and never updated for this layer’s signature, which takes packed variable-length inputs (cu_seqlens_, max_seqlen_) and rotary embeddings instead of attention_mask/layer_head_mask. This documents the parameters…
- 1c3633a #52127 — [CI/Build][CPU] Shrink triton-cpu-build layer by dropping build artifacts (#52127)
- 作者: Li, Jiang | +3/-0 | 1 个文件
- The vllm-triton-cpu-build stage in docker/Dockerfile.cpu clones and builds triton-lang/triton-cpu, but only the resulting wheel is needed by later stages/the final image. - The cloned triton-cpu source/build tree and Triton’s downloaded LLVM/MLIR toolchain (under /root/.triton) were being retained in the final layer, growing the stage to 3.46GB. - This PR removes the triton-cpu source tree after…
- 10bcad2 #52123 — Update CODEOWNERS (#52123)
- 作者: Cyrus Leung | +2/-0 | 1 个文件
As discussed offline with @Isotr0py ## Test Result —
- 7bbbf7c #51251 — [Core] Configure custom encoder cache managers from VllmConfig (#51251)
- 作者: hotTea | +36/-8 | 4 个文件
Expose custom encoder cache manager configuration through VllmConfig for both online and offline inference. Custom encoder cache managers may require policy-specific parameters in addition to the encoder cache size. This PR provides a generic configuration path while preserving compatibility with existing built-in and constructor-only cache managers. This PR: - Adds an opaque manager_config field …
- 373592e #52092 — [CPU] Ship triton-cpu wheel and fix several hardcoded pin_memory=True (#52092)
- 作者: Li, Jiang | +50/-41 | 8 个文件
- Build and install a pre-built triton-cpu wheel in the CPU build/test images instead of pip install-ing it from source inside CI, unblocking the Triton topk-topp kernel to run as a normal (non-soft-fail) test. - Move the topk-topp Triton kernel test out of the soft-fail CPU-ModelRunnerV2 Tests suite into CPU-Kernel Tests, and the linear-attention chunked-prefill correctness test into CPU-Language…
- 89c8401 #52037 — [Model] Skip unused Jina V5 output layers (#52037)
- 作者: kiroxu | +21/-3 | 1 个文件
Jina Embeddings V5 models are pooling-only, but their vLLM wrappers inherit causal-LM classes. Because these wrappers already declare themselves as pooling models, they bypass the generic pooling adapter that replaces generation-only output layers. The encoder/nano variant therefore retained an unused ParallelLMHead with shape [128256, 768]. This change applies the existing pooling-model no_init_w…
- b908a21 #48215 — [Model][LoRA] Add tower/connector LoRA support for Ultravox (#48215)
- 作者: Yangyi Gao | +345/-129 | 2 个文件
Purpose Part of #31479: enable enable_tower_connector_lora for Ultravox. This is not duplicate work. The open-PR and issue checks for #31479 show this as the only Ultravox tower/connector LoRA implementation; the other open PRs cover different multimodal model families. Ultravox previously could not apply LoRA to its audio tower and connector: - The tower/projector path was not fully built from …
- 5658391 #51882 — Remove NIXL reinstall step (#51882)
- 作者: ovidiusm | +1/-6 | 2 个文件
Remove the NIXL wheel reinstall step. It is not longer needed (NIXL fixed the dependency issues upstream) and it is not working correctly since it does not use the version pin. Currently it causes nixl-cu13 to float to 1.3.2 while nixl/nixl-cu12 stays on the kv_connectors.txt pin (nixl == 1.3.1). Verified by simulating the Dockerfile install path in a clean venv. ## Test Result Old path (nixl=…
- 903d2ef #51772 — [Attention][MLA] Fuse Kimi-K3 chunked-context K/V packing (#51772)
- 作者: Yongye Zhu | +996/-35 | 15 个文件
The K3 MLA layer delegated chunked-context prefill to impl._compute_prefill_context, which per chunk casts kv_nope to fp8, casts k_pe, concatenates [k_nope | k_pe], and re-quantizes a query the fused new-token epilogue had already quantized. This gives the layer its own context loop so that tail collapses into one kernel per chunk: fused_kimi_k3_mla_kv_concat{,_quant_fp8} reads the strided kv_b_pr…
- 6334491 #50513 — [XPU] update UMD to 26.27 (#50513)
- 作者: Yan Ma | +8/-8 | 1 个文件
Test Result —
- f9538af #52076 — [Core] Clearer comments in
BlockPool.free_blocks()(#52076)- 作者: Nick Hill | +10/-9 | 1 个文件
The comments explaining eviction precedence in BlockPool.free_blocks() were ambiguous/confusing. Make them clearer / more explicit.
- 9a276d6 #52003 — [Mypy Fix] Mypy fix for “vllm/model_executor/models/[cC][dD]” (#52003)
- 作者: Wentao Ye | +153/-66 | 24 个文件
Mypy fix for “vllm/model_executor/models/[cC][dD]”
- f9a0f62 #51879 — [KV Offload] Expose data-parallel topology to offloading backends (#51879)
- 作者: Ziqi Fan | +23/-1 | 8 个文件
Native KV-offloading backends currently receive the engine’s data_parallel_index, but not the total number of data-parallel replicas or the process-local DP rank. Consequently, OffloadingParallelConfig does not contain enough information to describe the DP topology. Add data_parallel_size and data_parallel_rank_local to OffloadingParallelConfig and populate them from ParallelConfig. data_parallel_…
- 6e502b6 #51813 — fix and test EPLB balancedness calculation (#51813)
- 作者: Julien Debache | +31/-3 | 2 个文件
Fix EPLB balancedness logging to aggregate rank load within each MoE layer. The current reduction uses the layer axis, so it can report perfect balance when one EP rank receives all tokens in every layer. ## Test Result The new regression case has equal token totals per layer with all tokens routed to one rank. Before the fix, it produced avg_tokens=100 and max_tokens=100; after the fix, it produc…
- 61826c1 #51624 — [Hardware][Power] Unqualized MoE Backend for Power (VSX) (#51624)
- 作者: Akash kaothalkar | +500/-6 | 6 个文件
This PR adds PowerPC specific unquantized backend support for fused MoE using Power10 VSX MMA instructions. Currently, grouped GEMM is not supported for Power architecture in vLLM. This PR introduces a Power/VSX specific unquantized CPU backend for Fused MoE. Key features include: - Addition of csrc/cpu/micro_gemm/cpu_micro_gemm_vsx.hpp to support grouped GEMM on Power architectures. - Use of Powe…
- 23f360e #52009 — [CI Bug] Fix ci moe test (#52009)
- 作者: Wentao Ye | +1/-1 | 1 个文件
Fixes https://buildkite.com/vllm/ci/builds/83443#019ff2a1-641e-4d2b-bca9-9eda8a060573 There are two errors here: 1. vLLM side, we name it triton test, but actually running the flashinfer path, this PR fixes the issue 2. the root cause of flashinfer is a bug upstream with TRT-LLM BF16 MoE, we may wait for their fix, not related to this PR Covered in CI
- 025d56a #52035 — [Build] Update DeepGEMM pin to deepseek-ai nv_dev tip (#52035)
- 作者: Yongye Zhu | +6/-6 | 2 个文件
Update both DeepGEMM pins (cmake/external_projects/deepgemm.cmake and tools/install_deepgemm.sh, which are documented to stay in sync) from the vllm-project/DeepGEMM fork at e21c821f to upstream deepseek-ai/DeepGEMM at 8b1392b978f5a03c828dd1711090d7fb50958b8a, the current tip of the nv_dev branch. The fork pin was kept because the plain nv_dev branch previously lacked SiTU support (see the removed…
- 7f7a32c #47808 — [Spec Decode] DSpark confidence-scheduled verification (#47808)
- 作者: Lucas Wilkinson | +1643/-111 | 40 个文件
Adaptively sizes the DSpark draft-verification budget from per-request confidence instead of always verifying every drafted token. Motivation: fixed-k speculation collapses at high concurrency — once the GPU saturates, verifying 7 drafts per request burns more compute than the accepted tokens return, dropping below non-speculative decoding (see table). ## Design - A Triton kernel ranks draft s…
- fe889ac #51311 — [K3 Perf] Flash kda out kernel for prefill, 1.1~1.4x kernel performance improvement (#51311)
- 作者: Wentao Ye | +36/-15 | 1 个文件
Use workspace manager to preallocate the memory for _flashkda_prefill, avoid re-allocate each time when we call the kernel. Acc covered in unit tests Perf can be seen in this AI generated script And we can get
- 0fb168e #52007 — [CI Bug] Fix ci qwen3.5 (#52007)
- 作者: Wentao Ye | +2/-2 | 2 个文件
Forward fix for https://github.com/vllm-project/vllm/pull/51908 Fixes https://buildkite.com/vllm/ci/builds/83443#019ff660-8010-40e8-851d-bb6479f24c68 Covered in CI
- b745d08 #51860 — [ROCm][K3] Dequantize the fp8 decode query for MLA backends without quant-query support - TRITON_MLA (#51860)
- 作者: Hongxia Yang | +7/-5 | 1 个文件
When enabling fp8 kv-cache dtype in DSpark speculative decoding, we got an assert error. THis PR is to put out a minimal fix to unblock this path. On ROCm, the DSpark draft auto-selects TRITON_MLA, the only backend supporting its non-causal multi-token blocks atm. TRITON_MLA dequantizes fp8 KV on load and takes a bf16 query, so Kimi-K3’s _decode_concat_cache assert on supports_quant_query_input ma…
- 324f452 #51464 — [ROCm] update triton in base docker for gluon compatibility (#51464)
- 作者: Hongxia Yang | +1/-1 | 1 个文件
To pick up https://github.com/ROCm/triton/pull/960 for fixing the gluon mla kernel compilation problem related to DistributedLinearLayout was rejected in the previous triton commit. With this, the gluon kernel can compile and serve K3 with accuracy. Context: Initially vllm/vllm-openai-rocm:nightly docker failed at the first MLA decode with a Gluon frontend error inside aiter/ops/triton/gluon/mla_g…
- 8151f2a #51999 — [Docs] Warn that –api-key does not gate all endpoints (#51999)
- 作者: Russell Bryant | +20/-1 | 2 个文件
The –api-key boundary is currently only documented in docs/usage/security.md: it authenticates the /v1, /v2, and /inference path prefixes, but other endpoints on the same HTTP server — notably /invocations, which exposes the same inference capabilities as /v1 — remain unauthenticated. Users (and LLM assistants, which frequently hit rate limits fetching the security docs) can easily miss this and …
- 9035151 #51255 — [Model] Add native Dots3 NOTE multimodal support (#51255)
- 作者: 范裕达 | +6468/-7 | 28 个文件
Add native vLLM support for the complete Dots3 NOTE model, including: - Text generation - Image understanding - Audio understanding - Native video processing with interleaved visual and audio inputs - FP8 MoE inference - MTP speculative decoding - OpenAI-compatible tool calling Dots3 NOTE uses one unified Hugging Face model type and architecture: - model_type: dots3_note - architectures: Dots3Note…
- b1b7520 #51841 — Avoid long-blocking H2D copies in ViT (#51841)
- 作者: Max Hu | +23/-5 | 2 个文件
Two host-to-device copies on the per-iteration critical path can take a very long time to return. Despite non_blocking=True, the cudaMemcpyAsync call blocks the calling thread, and the time it takes tracks the amount of GPU work that has already been queued. That undoes the host run-ahead that asynchronous scheduling and CUDA graphs exist to build up. Neither call site looks suspicious at a glance…
- 7aa248f #51935 — [XPU][CI/Release][2/N] add triton shim in xpu requirements (#51935)
- 作者: Kunshang Ji | +10/-5 | 3 个文件
to support xpu wheel release, we add triton shim wheel (https://wheels.vllm.ai/xpu/triton/) to resolve triton-xpu and triton(nvidia triton, dependency brought by other lib like xgrammar) conflict issue. this PR is adding this shim wheel as dependency. a side effect is we can not follow previos uv pip compile xpu job with –torch-backend xpu, so we changed .pre-commit-config.yaml CI ## Test Result …
- 8c011da #50787 — [XPU] Route block-quantized FP8 weights to the W8A8 kernel (#50787)
- 作者: Chaojun Zhang | +11/-1 | 2 个文件
This PR fixes an issue following #43645 and routes block-quantized weights to the W8A8 kernel unconditionally (no opt-in flag –linear-backend xpu required), while keeping the existing opt-in behavior for per-tensor/per-channel weights. ## Test Result Without this PR: server fails to start, worker crashes during load_model: Root cause: block-quantized weights were still routed to CompressedTen…
- 4eef91c #51831 — [Model] Support R3 capture with DeepGEMM MegaMoE (#51831)
- 作者: aoshen02 | +82/-0 | 4 个文件
Add routed-experts (R3) capture support to the DeepGEMM MegaMoE path used by DeepSeek V4 and Kimi K3. ## Why The existing binder recognizes MoE layers implemented through MoERunner. DeepGEMM MegaMoE bypasses that runner, even though its shared expert implementation already receives logical topk_ids before EPLB maps them to physical replicas. Enabling R3 therefore fails during model initialization …
- 10b7766 #51913 — [Attention] Move context_lens_tensor compute into GDN prefill path (#51913)
- 作者: Xin Yang | +1/-1 | 1 个文件
Move context_lens_tensor = m.compute_num_computed_tokens() in GDNAttentionMetadataBuilder.build() from the top of the method into the if num_prefills > 0 branch because it’s only used in prefill. Since context_lens_tensor is not used in decode path, this change avoids the tensor computed and discarded in decode path. e2e output throughput improves 7% at batch size 1 on H200. ## Profiling Main: PR:…
🦀 Rust Frontend
- 7553aac #51906 — [Frontend] Add routed-experts prompt offset (#51906)
- 作者: aoshen02 | +114/-61 | 13 个文件
- Add routed_experts_prompt_start to OpenAI chat/completion requests and SamplingParams, allowing clients to omit an already-known prompt prefix from returned R3. - Centralize NumPy-to-base64 serialization used by existing R3 responses and document the int32 expert-ID representation. - Keep OpenAI streaming behavior unchanged: R3 remains supported only on existing non-streaming responses. ## Why t…
- 152c913 #52098 — [Frontend] Log output token IDs at DEBUG level (#52098)
- 作者: yang rui | +52/-25 | 3 个文件
Allow operators to keep human-readable generated output logs without emitting output token IDs at the default INFO level. Following maintainer feedback, this now mirrors the existing request-input logging split instead of adding a new CLI flag: - INFO keeps generated text and the finish reason. - DEBUG additionally logs output token IDs. - –max-log-len continues to truncate both output text and t…
🔩 Misc
- 015660d #52145 — [Misc] Add missing return type annotations in outputs.py (#52145)
- 作者: Vineeta Tiwari | +13/-7 | 1 个文件
Purpose Add missing return type annotations to from_base() static methods and PoolingRequestOutput.repr() in vllm/outputs.py. All from_base() static methods on EmbeddingOutput, ClassificationOutput, ScoringOutput, EmbeddingRequestOutput, ClassificationRequestOutput, and ScoringRequestOutput lacked return type annotations. PoolingRequestOutput.repr() was missing -> str, making it the on…
- 7bc1660 #51931 — [Misc] Use VLLMValidationError in pooling input validation (#51931)
- 作者: Frank | +83/-5 | 3 个文件
Part of #48227. Like #51753, this is an independent file-level Step 5 migration. Migrate four caller-caused validation errors in vllm/entrypoints/pooling/base/io_processor.py from raw ValueError to VLLMValidationError: - conflicting offline pooling tasks - untrusted request-level chat templates - mismatched prompt and pooling parameter counts - mismatched prompt and LoRA request counts This preser…
📖 Documentation
- 443fa5a #51611 — [Doc] Fix stale rejection_sample_method and synthetic_acceptance_rate (#51611)
- 作者: QWERQWERQWE86 | +3/-2 | 1 个文件
Fixes #51609 Sync the –speculative-config table in docs/features/speculative_decoding/README.md with the current code: 1. rejection_sample_method (line 87): strict, probabilistic, synthetic (default strict) -> standard, synthetic, block (default standard); probabilistic now belongs to draft_sample_method. See #40651. 2. synthetic_acceptance_rate (line 88): split into synthetic_acceptance_rates (l…
⚡ Performance
- f962616 #51862 — [ROCm][Perf] Kimi-K3 Remove prefill pipeline stall in chunk KDA (#51862)
- 作者: kliuae | +188/-43 | 5 个文件
On ROCm’s Kimi-K3 path, each prefill/mixed step has a stall in prepare_chunk_indices, caused by two interacting factors: - tolist() D2H copy makes the host wait until the device queue is emptied - Host ops executing before the next H2D copy, while the device is sitting idle The data that the host waits on is already sitting on the host. At each step, GDNAttentionMetadataBuilder builds chunk_indice…
- 3d204df #52024 — Revert “[Perf][ROCm] Dual-stream decode with hipgraphs” (#52024)
- 作者: Simon Danielsson | +55/-55 | 2 个文件
Reverts vllm-project/vllm#48223
✨ New Feature
- 50ba4bc #49577 — [Feature] Mask Replay (#49577)
- 作者: vx120 | +556/-5 | 24 个文件
This PR adds experimental support for sampling distribution replay. A sampling mask represents the vocabulary support retained after top-k/top-p filtering. It is not an attention mask and does not affect causal attention or KV-cache behavior. When enabled, vLLM returns the sampling support for each generated token in a CSR-style representation: The feature is opt-in and does not change default…
🧪 CI/Tests
- 8eb35c5 #52064 — [CI] Mirror external test assets in vLLM S3 (#52064)
- 作者: Kevin H. Luu | +30/-41 | 5 个文件
- mirror externally hosted video, image, and GSM8K test assets into the public vLLM S3 bucket - update affected tests to use the shared VLLM_S3_BUCKET_URL constant ## Why Buildkite build 83608 had multiple test failures caused by direct dependencies on third-party hosts, including RemoteDisconnected while fetching OpenCV and Bogotobogo video fixtures. The mirrored objects are publicly readable fro…
- caf9e8f #52043 — [CI] Force source builds for hybrid dependencies (#52043)
- 作者: Andreas Karatzas | +14/-14 | 2 个文件
Hybrid language-model CI began failing before pytest while mamba-ssm and causal-conv1d probed guessed GitHub release-wheel URLs. The requested CUDA/Torch and ROCm/Torch wheels do not exist, so the normal path is a caught HTTP 404 followed by a source build. During today’s intermittent GitHub connectivity problems, some requests instead ended with Remote end closed connection without response, whic…
- 6accb77 #51911 — [CI] Add registry layer cache to x86 CPU image build (#51911)
- 作者: Li, Jiang | +172/-24 | 2 个文件
The x86 CPU CI image build (.buildkite/image_build/image_build_cpu.sh) used classic docker build with no layer cache, so every build recompiled csrc/rust from scratch. This mirrors the registry-based BuildKit layer cache already used for the CUDA CI image build (image_build.sh). ## What changed - Switch from docker build + docker push to docker buildx build –push using a dedicated vllm-cpu-builde…