共 76 个 commit,涉及 381 个文件,+23699/-3694 行变动。
概要
| 统计项 | 数值 |
|---|---|
| Commit 数 | 76 |
| 变更文件 | 381 |
| 新增行数 | +23699 |
| 删除行数 | -3694 |
Commit 列表
📦 Other
- 8151f2a #51999 — [Docs] Warn that –api-key does not gate all endpoints (#51999)
- 作者: Russell Bryant | +20/-1 | 2 个文件
The –api-key boundary is currently only documented in docs/usage/security.md: it authenticates the /v1, /v2, and /inference path prefixes, but other endpoints on the same HTTP server — notably /invocations, which exposes the same inference capabilities as /v1 — remain unauthenticated. Users (and LLM assistants, which frequently hit rate limits fetching the security docs) can easily miss this and …
- 9035151 #51255 — [Model] Add native Dots3 NOTE multimodal support (#51255)
- 作者: 范裕达 | +6468/-7 | 28 个文件
Add native vLLM support for the complete Dots3 NOTE model, including: - Text generation - Image understanding - Audio understanding - Native video processing with interleaved visual and audio inputs - FP8 MoE inference - MTP speculative decoding - OpenAI-compatible tool calling Dots3 NOTE uses one unified Hugging Face model type and architecture: - model_type: dots3_note - architectures: Dots3Note…
- b1b7520 #51841 — Avoid long-blocking H2D copies in ViT (#51841)
- 作者: Max Hu | +23/-5 | 2 个文件
Two host-to-device copies on the per-iteration critical path can take a very long time to return. Despite non_blocking=True, the cudaMemcpyAsync call blocks the calling thread, and the time it takes tracks the amount of GPU work that has already been queued. That undoes the host run-ahead that asynchronous scheduling and CUDA graphs exist to build up. Neither call site looks suspicious at a glance…
- 7aa248f #51935 — [XPU][CI/Release][2/N] add triton shim in xpu requirements (#51935)
- 作者: Kunshang Ji | +10/-5 | 3 个文件
to support xpu wheel release, we add triton shim wheel (https://wheels.vllm.ai/xpu/triton/) to resolve triton-xpu and triton(nvidia triton, dependency brought by other lib like xgrammar) conflict issue. this PR is adding this shim wheel as dependency. a side effect is we can not follow previos uv pip compile xpu job with –torch-backend xpu, so we changed .pre-commit-config.yaml CI ## Test Result …
- 8c011da #50787 — [XPU] Route block-quantized FP8 weights to the W8A8 kernel (#50787)
- 作者: Chaojun Zhang | +11/-1 | 2 个文件
This PR fixes an issue following #43645 and routes block-quantized weights to the W8A8 kernel unconditionally (no opt-in flag –linear-backend xpu required), while keeping the existing opt-in behavior for per-tensor/per-channel weights. ## Test Result Without this PR: server fails to start, worker crashes during load_model: Root cause: block-quantized weights were still routed to CompressedTen…
- 4eef91c #51831 — [Model] Support R3 capture with DeepGEMM MegaMoE (#51831)
- 作者: aoshen02 | +82/-0 | 4 个文件
Add routed-experts (R3) capture support to the DeepGEMM MegaMoE path used by DeepSeek V4 and Kimi K3. ## Why The existing binder recognizes MoE layers implemented through MoERunner. DeepGEMM MegaMoE bypasses that runner, even though its shared expert implementation already receives logical topk_ids before EPLB maps them to physical replicas. Enabling R3 therefore fails during model initialization …
- 10b7766 #51913 — [Attention] Move context_lens_tensor compute into GDN prefill path (#51913)
- 作者: Xin Yang | +1/-1 | 1 个文件
Move context_lens_tensor = m.compute_num_computed_tokens() in GDNAttentionMetadataBuilder.build() from the top of the method into the if num_prefills > 0 branch because it’s only used in prefill. Since context_lens_tensor is not used in decode path, this change avoids the tensor computed and discarded in decode path. e2e output throughput improves 7% at batch size 1 on H200. ## Profiling Main: PR:…
- 10f9b5d #51923 — [CI/Release][XPU] fix workdir path for triton shim job (#51923)
- 作者: Kunshang Ji | +4/-1 | 1 个文件
fix https://buildkite.com/vllm/release-v2/builds/5061#019ff414-8a41-4375-9f62-df100fe5b6ef ## Test Result —
- f4e9bb0 #51759 — [CI/Release][1/N][XPU] Publish XPU Triton shim index (#51759)
- 作者: Kunshang Ji | +95/-0 | 2 个文件
Summary This PR adds a release-pipeline job to publish the Intel XPU Triton compatibility shim to s3://vllm-wheels/xpu/. - Downloads the pinned triton-3.7.2+xpu wheel from the Intel XPU Triton release. - Verifies the wheel against its fixed SHA256 checksum before publishing. - Generates and uploads PEP 503 indexes and metadata for: - https://wheels.vllm.ai/xpu/ - https://wheels.vllm.ai/xpu/trito…
- 3fa8922 #51652 — [CI/Build] Use file rendezvous for local distributed tests (#51652)
- 作者: Stellar鱼 | +31/-12 | 3 个文件
Continue the test-only follow-up identified in #51275 by removing probe-then-bind TCP rendezvous from two single-node distributed test launchers. get_open_port() releases its probe socket before the spawned workers initialize their real process group. Under parallel CI, another process can claim that port during the gap and make an otherwise-correct test fail with EADDRINUSE. This PR: - switches t…
- 20727e2 #50268 — [Hardware][AMD] Enable fused bf16→fp32 router GEMM on ROCm (#50268)
- 作者: Matvei Pashkovskii | +112/-3 | 2 个文件
This pull request addresses issue #50267 by enabling fused bf16→fp32 GEMM operations on AMD ROCm hardware for MoE router gates. ## Problem Statement The MoE router gate requires out_dtype=torch.float32 for the grouped_topk operation. However, on ROCm, the fused fp32-output GEMM implementation was gated exclusively to CUDA platforms. This caused a fallback to bf16 GEMM followed by a separate cast o…
- 793ca69 #51729 — [Docs][RL] Rewrite weight-transfer docs; standardize examples (#51729)
- 作者: Aaron Hao | +1226/-1157 | 16 个文件
Closes out the two follow-ups that PR 3 of the trainer-side weight-transfer split (#48042, #48981, and the NCCL/sparse PR) deferred, and converges the RL examples on a single inference-side pattern. Docs (docs/training/weight_transfer/) - Every trainer-side snippet documented the static trainer_send_weights / *TrainerSendWeightsArgs path that PR 2 and PR 3 removed, so all four pages are rewritten …
- 9b33826 #51878 — [Tools] vLLM Recipes conversion : support different data types variants and strategies (#51878)
- 作者: Louie Tsai | +167/-25 | 2 个文件
Support vLLM recipes conversion for Different data types and multi or single nodes strategies manual test ## Test Result passed. —
- 466855a #47017 — [ROCm] Enable DeepSeek-V4 on gfx11 (#47017)
- 作者: PikaPikachu | +24/-2 | 3 个文件
This PR enables DeepSeek-V4 checkpoints on ROCm gfx11/RDNA devices. It removes Python-side blockers in the ROCm sparse-indexer path and allows DeepSeek-V4 checkpoints mapped to INCConfig to pass ROCm platform validation: - Register ROCm sparse-indexer ops on gfx11. - Let SparseAttnIndexer.forward_hip() delegate to the ROCm sparse-indexer custom op and use the op-level fallback instead of failing e…
- 1d2d83a #49718 — [Attention] Add FlashInfer XQA decode support on SM12x (#49718)
- 作者: Andrii Skliar | +413/-57 | 4 个文件
Enable FlashInfer XQA decode on SM120/SM121 through FlashInfer’s dedicated XQA API. - Select XQA for SM12x decode. - Support uniform and ragged speculative decode through q_len_per_req and q_cu_seq_lens. - Build the packed causal or non-causal draft mask required by XQA. - Support sliding-window decode, attention sinks, non-causal draft attention, and uniform-batch CUDA graphs on SM12x. - Keep the…
- a367cdb #48789 — [Profiler] Add minimal Triton Proton profiling backend (#48789)
- 作者: 鐘天楽 | +638/-21 | 7 个文件
Add Triton Proton as an optional worker profiling backend alongside PyTorch and CUDA profiling. This first PR deliberately supports NVIDIA GPUs through CUPTI and graph-free execution only, keeping the core worker integration small and reviewable. - Expose typed Proton context, data, CUPTI backend, output format, mode, and Triton hook configuration. - Keep Proton lazily imported so other profiler u…
- 61874f9 #51612 — [4/N][KV-Cache Layout Refactor] Promote local KV cache specs via a class-changing replace helper (#51612)
- 作者: Lucas Wilkinson | +35/-36 | 2 个文件
_promote_local_kv_cache_specs rebuilds each promoted spec with a hand-written constructor call per class, and the explicit field lists have drifted: the MLA and chunked-local promotions silently drop kv_quant_mode (mis-sizing promoted pages for quantized KV caches), and the chunked-local promotion drops head_size_v. This PR adds replace_as() — dataclasses.replace generalized to rebuild a spec as a…
- cb30f6f #51736 — [Testing] Fix test_sharded_state_loader (#51736)
- 作者: Richard Zou | +4/-4 | 1 个文件
This test compares the following: 1. load checkpoint -> generate outputs from prompts 2. load checkpoint -> reshard checkpoint -> reload checkpoint -> generate output from prompts Generating the output from prompts assumes batch invariance; if the vLLM scheduler schedules a different amount of batches across 1 and 2 then we can get different results. This PR forces the generation to do one sequenc…
- 0f0cb91 #49758 — [ROCm][MoE] Fix expert_map vs AITER expert_mask for non-AITER experts under EP (#49758)
- 作者: Rohan Potdar | +70/-59 | 6 个文件
RoutedExperts.expert_map returned the AITER 0/1 expert_mask whenever AITER fused MoE is enabled (VLLM_ROCM_USE_AITER=1), regardless of which experts kernel actually runs. Non-AITER kernels (Triton moe_wna16, MXFP4 emulation, Marlin, …) expect the canonical global→local expert_map (-1 for non-local) and misread the 0/1 mask as slot indices → every token collapses onto local slot 0/1 → garbage under…
- b2ab096 #51857 — [Docs] Fix broken autorefs cross-reference in TurboQuant v2 docstring (#51857)
- 作者: Harry Mellor | +1/-1 | 1 个文件
The docs build emits exactly one WARNING, from the triton_turboquant_decode_v2 module docstring: The docstring writes pair_table[i][j] unquoted. Markdown parses [i][j] as a full reference link (text i, reference j), mkdocs-autorefs then tries to resolve j as a cross-reference target and fails. Wrapping the expression in backticks makes it render as code, which is also how the same construct is alr…
- 1ab2801 #50907 — [ROCm] Remove stale SDPA and skinny GEMM workarounds (#50907)
- 作者: Andreas Karatzas | +500/-558 | 56 个文件
- Remove ROCm-wide Math SDPA forcing from Transformers multimodal execution and affected test conftests. - Remove blanket skinny GEMM disable overrides while retaining independent backend and determinism settings. - Make non-contiguous skinny GEMM activations contiguous and fall back safely for unsupported weight or bias layouts. - Use default speech attention on MI250 while retaining supported AI…
- 52be12c #50569 — feat: allow shared expert overlapping for FlashInfer one-sided all-to-all (#50569)
- 作者: Julien Debache | +9/-2 | 1 个文件
Allow shared expert overlap when using EPLB with the flashinfer_nvlink_one_sided all2all backend. Previously, PR #28377 disabled shared expert overlap for all non-allgather_reducescatter backends under EPLB due to correctness issues observed with deepep_low_latency. This PR exempts flashinfer_nvlink_one_sided from that restriction after verifying it produces correct results with overlap enabled. S…
- 9cc347a #45042 — Enable
MiniCPMVfor vLLM in CI (#45042)- 作者: Harry Mellor | +8/-0 | 1 个文件
Restores the max_transformers_version cap on the MiniCPMV test registry entry, scoped to the HF runner only (transformers_version_reason={“hf”: …}). This PR previously flipped an existing “vllm” reason to “hf”. Since then #48413 removed the cap outright, so it has been rewritten: the cap is re-added, with the correct reason. ## Why the cap is needed again Nightly build 83094 (Multi-Modal Models …
- c65aa2e #51668 — Bump Transformers version to 5.15.0 (#51668)
- 作者: Harry Mellor | +23/-17 | 11 个文件
- 4988df2 #51726 — [Config] Update default
_max_num_batched_tokensfrom 8192 to 16384 (#51726)- 作者: Wentao Ye | +12/-31 | 2 个文件
Update this num to a larger num when gpu memory is enough, perf can be seen https://github.com/vllm-project/vllm/pull/51725. Note that SGLang use the same num I make the PRs separate so easier to review and revert if any issue bumps up.
- dd7cc85 #51407 — Add MoE output contract for MoE tail fusion (#51407)
- 作者: Jee Jee Li | +127/-13 | 4 个文件
Test Result —
- b64a270 #51447 — Bound generation inputs before expensive work (#51447)
- 作者: Clinton Thomas | +421/-37 | 11 个文件
Bound generation inputs before expensive work ## What this fixes Five ordinary request values could make vLLM perform work far larger than the JSON body before an effective limit ran. Examples included tokenizing 50,000 bad words before a later rejection, scanning 500,000 stop strings after every generated token, and scanning an entire DeepSeek message history again for every message. Before thi…
- 5b184f7 #51308 — connects vLLM Recipes with vLLM’s native config-based deployment and benchmark (#51308)
- 作者: Louie Tsai | +773/-43 | 6 个文件
This PR connects vLLM Recipes with vLLM’s native config-based deployment flow. It adds a tool that converts a hardware-specific Recipe into: * config.yml — used by vllm serve –config * env.sh — Recipe-specific environment variables The docs are updated to make this workflow discoverable from CPU model support, deployment, benchmarking, and server configuration pages. ### Flow ### Example See …
- f863387 #51770 — [XPU] Fix UVA weight offloading (non-pinned-tensor views and static Triton launcher) (#51770)
- 作者: Chaojun Zhang | +33/-1 | 2 个文件
Fix two startup crashes when using weight offloading (–cpu-offload-gb) on XPU. ### Issue 1: Pinned-memory assertion failure The assertion fails for zero‑element tensors (common in quantized checkpoints) or when pinning is disabled via VLLM_WEIGHT_OFFLOADING_DISABLE_PIN_MEMORY=1. The underlying XPU kernel already handles both cases, but the Python wrapper enforces the assertion prematurely. Fix: D…
- d8f8400 #51780 — [Model] Enable tower and connector LoRA for Keye (#51780)
- 作者: liushujia122 | +8/-0 | 1 个文件
- Implement get_num_mm_encoder_tokens() and get_num_mm_connector_tokens() on BaseKeyeModule. - Convert between the post-merge multimodal token count seen by the language model and the pre-merge activation length used by the vision tower and Keye projector. Part of #31479. ## Implementation Keye reports language-model image tokens after spatial merging: The vision tower operates on the unmerged pat…
- 490259c #50977 — profiler: add PrivateUse1 activity support for custom backends (#50977)
- 作者: David Holtz | +2/-1 | 1 个文件
This change enables profiler support of vLLM for custom or third-party backends using PyTorch’s PrivateUse1 device / activity namespace. For example, once integrated in spyre-inference this enables profiling on the IBM’s spyre device. n/a ## Test Result n/a —
- 52c70b2 #50826 — [XPU] [Linear] enable torch linear backend for blockwise gemm on xpu (#50826)
- 作者: zofia | +2/-0 | 2 个文件
This pr enables the torch backend on xpu.
- a311916 #46849 — [MRV2][Spec] Fuse AR speculator multi-step decodes back into one CUDA graph (#46849)
- 作者: Yizhou | +492/-49 | 10 个文件
This PR restores fused multi-step CUDA graph execution for autoregressive speculative decoding in Model Runner V2. #41162 fixed stale attention metadata by rebuilding it and replaying a separate CUDA graph for every draft step. While correct, that design reintroduces per-step Python dispatch, metadata construction, and CUDA graph launch overhead. This PR instead captures the post-prefill draft loo…
🐛 Bug Fix
- 8e95890 #51843 — [Bugfix] Disable fine-grained prefix-cache hits for incompatible hybrid KV layouts (#51843)
- 作者: Michael Goin | +96/-3 | 3 个文件
Fine-grained prefix-cache hits are enabled for hybrid models containing Mamba align groups. However other groups, such as a sliding-window DSpark drafter, may use KV cache managers that only support block-aligned lookups. This previously caused an assertion when prefix caching was enabled. This change: - Enables fine-grained hits only when every KV cache manager supports the required lookup granul…
- 4c51ceb #51139 — [Bugfix][Multimodal] Invalidate retained PyNvVideoCodec decoder after failure (#51139)
- 作者: dmai-afk | +134/-1 | 3 个文件
Title [Bugfix][Multimodal] Invalidate retained PyNvVideoCodec decoder after failure # Body Fixes #51138 - invalidate a retained PyNvVideoCodec decoder slot when a borrowed operation exits with an error - clear stale slot state before constructing a replacement decoder - publish the replacement decoder only after construction succeeds - add CPU-only regression coverage for borrow-time invalidatio…
- 86e2ab5 #51120 — [Bugfix][Frontend] Return 400 for invalid PyNvVideoCodec video input (#51120)
- 作者: dmai-afk | +110/-5 | 3 个文件
- translate PyNvVideoCodec.PyNvVCException raised while opening malformed video into a sanitized ValueError - rely on vLLM’s existing ValueError handler to return BadRequestError/HTTP 400 instead of HTTP 500 - preserve the native exception as the cause and leave unrelated decoder exceptions unchanged ## Root cause A malformed user-supplied video reaches PyNvVideoCodec.SimpleDecoder during multimod…
- e60f3c4 #46845 — [Bugfix] Fix MiniMax-M3 compressed-tensors FP8 MoE SwiGLU params (#46845)
- 作者: Tan Pin Siang | +34/-0 | 2 个文件
This PR fixes MiniMax-M3 accuracy when served from a compressed-tensors W8A8 FP8 MoE checkpoint. MiniMax-M3 uses hidden_act=“swigluoai” with non-default SwiGLU parameters: - swiglu_alpha=1.702 - swiglu_beta=1.0 - swiglu_limit=7.0 The compressed-tensors FP8 MoE path already forwarded swiglu_limit, but did not forward swiglu_alpha or swiglu_beta into make_fp8_moe_quant_config. As a result, MiniMax-M…
- aeece10 #51928 — [Bugfix][XPU] Run GDN attention as eager break under breakable cudagraph (#51928)
- 作者: Huanxing | +2/-1 | 1 个文件
When breakable graph, Qwen3.5 on XPU produced garbage output. This PR is fix the GDN (Gated Delta Net) recurrent attention on XPU. The XPU GDN attention code reads and writes the conv and ssm state caches in place using data-dependent metadata. Under breakable-cudagraph capture the capture-time state addressing is frozen into the graph, so replay corrupts the recurrent state and hybrid models (e.g…
- 3ac9525 #50074 — [Bugfix][Quantization] Reuse online NVFP4 MoE kernel across reloads (#50074)
- 作者: Matej Sirovatka | +60/-10 | 2 个文件
Online NVFP4 MoE models re-run process_weights_after_loading whenever fresh BF16 weights are loaded and quantized in place. _setup_kernel currently also replaces the MoE kernel and quantization-config objects on every reload. Compiled and captured execution paths still reference the original kernel object, so replacing it can leave reloads executing against stale state and eventually produce non-f…
- f2906b4 #49505 — [Bugfix] Avoid repeated layerwise reload warning scans (#49505)
- 作者: aoshen02 | +33/-2 | 2 个文件
Avoid repeatedly scanning all concurrently buffered layerwise-reload state for every incoming weight tensor. The overlap warning is now evaluated only when a layer first enters LOADING_LAYERS, and its detailed size/name calculation runs only when the set first transitions from one loading layer to two. Related to #48312. ## Problem online_process_loader runs once for every weight or packed shard a…
- 3962042 #46747 — [Bugfix][V1][Multimodal] Recover from P0/P1 processor cache drift (#46747) (#46747)
- 作者: WillZZZy | +236/-113 | 13 个文件
Summary: The multi-modal IPC processor cache keeps two copies that are assumed to stay mirrored: a metadata-only “shadow” on P0 (the API / front-end process) and the real tensor cache on P1 (the engine process). On a P0 shadow hit, P0 forwards data=None and trusts that P1 still holds the item. Because the two processes update their byte-budgeted LRU caches in different orders — P0 processes multi-…
- d462dee #51872 — [Bugfix][Triton] Make fp8_min/fp8_max constexpr in _quantize_pad_fp8_kernel (#51872)
- 作者: Max Hu | +2/-2 | 1 个文件
_quantize_pad_fp8_kernel in vllm/kernels/triton/qkv_padded_fp8_quant.py fails to compile under torch.compile when the ViT encoder is compiled with FP8 encoder attention (compile_mm_encoder), because its fp8_min / fp8_max scalar arguments are typed as fp64 by Inductor. This PR declares them tl.constexpr so the kernel composes with torch.compile. ## Error When the FP8 ViT encoder attention path is c…
- f067737 #51840 — [Bugfix][TieredOffloading] : Return HIT_PENDING when KV promotion is triggered (#51840)
- 作者: Varun Sundar Rabindranath | +15/-15 | 2 个文件
fixes https://github.com/vllm-project/vllm/issues/51439 ## Test Result
- 0a94d85 #51865 — [Bugfix][MRV2] Require all requests to be decoding for uniform-decode dispatch (#51865)
- 作者: Nick Hill | +341/-54 | 9 个文件
This is a copy of https://github.com/vllm-project/vllm/pull/50532 from @rchalamala with some additional rework to avoid redundant computation. Part of the prepare_inputs logic is split into a separate gather_batch_req_state method which runs before the dp token count / cuda-graph mode synchronization.
- 3e372c5 #51837 — [Bugfix][ROCm] Give KV-first attention blocks their own page in hybrid models (#51837)
- 作者: stefankoncarevic | +177/-1 | 2 个文件
On MI300, DSpark speculative decoding on RedHatAI/Qwen3.6-35B-A3B-NVFP4 produces garbage for most requests in a batch. GSM8K accuracy drops to 0.20 against 0.95 for the same target model without speculation, and test_dspark_correctness_and_acceptance_rate[qwen3.6-speculators] fails. The same test passes on H200. Hybrid models share a single page pool between the full-attention layers, the Mamba/GD…
- 36f4630 #50727 — [Bugfix][MoE] Fix fused block-scale orientation (#50727)
- 作者: Andreas Karatzas | +80/-29 | 5 个文件
PR #50137 fixed fused per-channel scale loading, but shape-based orientation inference cannot reliably recover the checkpoint layout after both dimensions have been divided by the quantization block size. This causes Qwen3-VL’s fused gate/up and down-projection block scales to load in the wrong orientation, producing this large-model acceptance failure. Represent the fused checkpoint layout explic…
- ded6c45 #51850 — [Bugfix] Support HF-config compat for Inkling (#51850)
- 作者: Michael Goin | +19/-1 | 1 个文件
Accept both OG Inkling and Transformers config.json field names for MoE/dense intermediate sizes and conv kernel size, so new quantized checkpoints load without rewriting Hub configs. See https://github.com/huggingface/transformers/blob/fe747d88a3296bd94d426db2717f232f9d4afdb7/src/transformers/models/inkling/configuration_inkling.py#L95 for reference ## Test Result —
- cc668c5 #51854 — [CI][Bugfix][V1] Remove stale FlashAttention metadata arguments (#51854)
- 作者: Andreas Karatzas | +0/-2 | 1 个文件
- Regression: #51756 removed FlashAttentionMetadata.sliding_window. - Detection: Buildkite main build #83388 first exposed two stale V1 test constructors; the V1 Core job in build #83390 reproduced the same failure. - This PR removes the obsolete sliding_window=None arguments from both speculator metadata fixtures. PR #51756 correctly moved sliding-window ownership out of FlashAttentionMetadata, b…
- 4f2f31b #51363 — [Bugfix][Attention] Forward per-head FP8 descales through FA4 (#51363)
- 作者: Yi Liu | +8/-0 | 1 个文件
Forward per-head FP8 Q/K/V descales from vLLM’s FlashAttention interface into the FA4 CuTe forward kernel. ## Root cause The FA3 path forwarded q_descale, k_descale, and v_descale, but the FA4 path dropped them. Per-head FP8 Q/K/V tensors were therefore consumed without their calibration scales, producing garbage output for llm-compressor attention/KV-quantized checkpoints. - Pass Q/K/V descales t…
- e3fe212 #51749 — [Bugfix] Generalize KV block zeroing to
AttentionSpec(#51749)- 作者: Michael Goin | +134/-10 | 4 个文件
- Record newly allocated AttentionSpec block IDs when worker-side KV zeroing is required. - Include all allocating attention groups in KVBlockZeroer. - Add focused sliding-window and chunked-local regressions. ## Root cause KVCacheConfig.needs_kv_cache_zeroing is enabled for hybrid Mamba and mixed-precision caches, but the scheduler only reported an exact allowlist of full/MLA block types and the …
- 513f83e #51756 — [Bugfix] Take the sliding window from the layer, not the KV cache group (#51756)
- 作者: Nick Hill | +116/-26 | 4 个文件
One KV cache group can hold both windowed and global layers — Gemma-3 with –disable-hybrid-kv-cache-manager promotes its sliding-window layers to full-attention storage, and the merged group spec still records the window. FlashAttention read that window off the group spec and applied it to every layer in the group, so the global layers silently lost everything older than the window; CPU attention…
- 457a5f3 #50020 — [Bugfix][MRV2] Support encoder timing stats in model runner V2 (#50020)
- 作者: Guan-Ming Chiu | +99/-22 | 7 个文件
- vllm bench mm-processor crashes on model runner V2: get_encoder_timing_stats only exists in V1 - Port it to V2: timing lives on EncoderRunner, gated on enable_mm_processor_stats; EncoderTimingStats moves to worker/utils.py for sharing - Not a duplicate: no open PR touches MRV2 encoder timing - pytest tests/v1/worker/test_encoder_runner.py — adds registry accumulate/clear test - vllm bench mm-pro…
- 5af7c8d #51812 — [Bugfix] Align Qwen GDN gates with speculative tokens (#51812)
- 作者: Jiangyun Zhu | +6/-2 | 1 个文件
Fix Qwen GDN speculative decoding when a mixed batch places non-speculative tokens before speculative tokens. mixed_qkv is gathered with spec_token_indx, but the fused recurrent update previously received the unsorted a and b gate tensors. The kernel consumes the first T_spec gate rows, so the gates could belong to different tokens than the gathered Q/K/V rows. Gather a and b with the same indices…
- 0fb9897 #51819 — [Bugfix][MoE] Support GELU tanh in FlashInfer B12x MoE (#51819)
- 作者: Andrii Skliar | +6/-1 | 1 个文件
Adds MoEActivation.GELU_TANH support to the FlashInfer B12x MoE backend so Gemma 4 NVFP4 can use –moe-backend flashinfer_b12x. ## Validation Validated with nvidia/Gemma-4-26B-A4B-NVFP4 on DGX Spark using FlashInfer 0.6.16.post3
- 87668ab #51768 — [Bugfix] Guard DeepSeek V4 MRV1 piecewise CUDA graphs (#51768)
- 作者: Woosuk Kwon | +73/-0 | 2 个文件
Default DeepseekV4ForCausalLM to Model Runner V2 and reject only the known-broken configuration: - DeepSeek V4 - Model Runner V1 - PIECEWISE or FULL_AND_PIECEWISE CUDA graphs The validation runs after CUDA-graph mode resolution. MRV1 therefore remains available with eager/NONE, FULL, or FULL_DECODE_ONLY; other model architectures are unaffected. ## Why #51430 exposed a correctness problem in the l…
- 12bea3e #51145 — [Bugfix][ROCm] Fix DeepSeek V4 DSpark probabilistic startup (#51145)
- 作者: Tuukka Sarvi | +2/-0 | 1 个文件
DeepSeek V4 DSpark uses a full-vocabulary draft model, so draft token ids are already target token ids and no draft-to-target remapping is needed. The CUDA implementation exposes this by defining draft_id_to_target_id = None, but the ROCm platform implementation was missing the same marker. The shared DSpark speculator reads model.draft_id_to_target_id when draft_sample_method=“probabilistic” to d…
- 78e7fdd #51627 — [Bugfix][CPU] Make the Apple Silicon BF16 probe fall back instead of raising (#51627)
- 作者: Kush Zingade | +8/-5 | 1 个文件
Problem In CpuPlatform.supported_dtypes, the macOS arm64 branch probes for BF16 with: macOS does not publish the hw.optional.arm.FEAT_BF16 OID when the CPU does not have the feature. sysctl then prints unknown oid to stderr and exits 1. check_output raises CalledProcessError on a nonzero exit, so it never returns a value other than b"1", and the return [torch.float16, torch.float32] line below …
- 0fec3d6 #51622 — [Bugfix][KV Offload] Centralize shared mmap cleanup in CPU worker (#51622)
- 作者: AlexHuang | +220/-13 | 2 个文件
Fix shared mmap ownership and shutdown ordering in the CPU KV offloading worker. CPUOffloadingWorker creates one set of mmap-backed CPU tensors and passes the same tensors to both the GPU-to-CPU and CPU-to-GPU handlers. Previously, the region cleanup responsibility was attached to one direction, even though both directions could still have asynchronous transfers using the same backing memory. This…
- ce07118 #51766 — [Bugfix][Core] Preserve Mamba running CoW after external hits (#51766)
- 作者: Dao007forever | +73/-0 | 2 个文件
Preserve running-request copy-on-write semantics for Mamba align requests whose externally loaded prefix already provides every block needed by their first continuation. The observed Kimi-K3 geometry is: ## Root cause allocate_external_computed_blocks() can populate the request’s Mamba block table, but when the first continuation stays in the same Mamba block, allocate_new_blocks() returns early w…
🧪 CI/Tests
- 6accb77 #51911 — [CI] Add registry layer cache to x86 CPU image build (#51911)
- 作者: Li, Jiang | +172/-24 | 2 个文件
The x86 CPU CI image build (.buildkite/image_build/image_build_cpu.sh) used classic docker build with no layer cache, so every build recompiled csrc/rust from scratch. This mirrors the registry-based BuildKit layer cache already used for the CUDA CI image build (image_build.sh). ## What changed - Switch from docker build + docker push to docker buildx build –push using a dedicated vllm-cpu-builde…
- a53ad85 #51905 — [XPU][CI]Change to use global VLLM_DISABLE_COMPILE_CACHE=1 in Intel GPU CI (#51905)
- 作者: xiangdong | +13/-17 | 9 个文件
Change to use global VLLM_DISABLE_COMPILE_CACHE=1 in Intel GPU CI ## Test Result —
- 23f25aa #50804 — [CI] Stabilize tensor IPC multiprocessing tests (#50804)
- 作者: Andreas Karatzas | +271/-176 | 1 个文件
- Use an explicit spawn context for all queues, events, barriers, and worker processes without changing the process-wide multiprocessing default. - Give spawned workers one shared 60-second startup and result deadline to accommodate slower ROCm imports. - Tag encoder and decoder results by role instead of assuming cross-process queue ordering. - Preserve the shorter tensor transport timeouts so a …
- 045a986 #51877 — [ROCm][CI] Speed Up ROCm Skinny GEMM Tests (reduced parameterizations, (#51877)
- 作者: Micah Williamson | +72/-35 | 2 个文件
This PR reduces the test_rocm_skinny_gemm suite from a ~2 hour runtime to about 10 seconds by reducing the number of parameterizations from 11040 to 2644, and removing the unnecessary environment cleanup between tests. We now just cleanup the environment one time at the end of the entire module rather than needlessly eating the 0.3s cost to cleanup after each parameterization.
- ca9c8cb #51832 — [CI] Support partial torch requirement contexts (#51832)
- 作者: Taneem Ibrahim | +1/-1 | 1 个文件
A clean CUDA CI image build runs use_existing_torch.py from a partial test-deps context without pyproject.toml, causing dependency setup to fail before tests. Treating that file as optional preserves the dependency cache while supporting partial contexts. ## Reproducer Run python use_existing_torch.py –prefix with requirements/ present and no pyproject.toml. ## Output on main / on branch Main: Fi…
- d648236 #51806 — [XPU][CI] Fix ExampleConnector KV cache device selection (#51806)
- 作者: liuzhenwei | +3/-3 | 3 个文件
- Load KV cache tensors onto the destination cache tensor’s device instead of hard-coding CUDA. - Add example connector test into XPU CI. ## Test Result —
🔧 Refactor
- 02ac178 #51917 — [Refactor][MRV2] Unify uniform decode token count helper (#51917)
- 作者: Lucas Wilkinson | +12/-24 | 3 个文件
Remove the redundant get_uniform_token_count wrapper and route its callers through get_uniform_decode_token_count. Dummy and capture-only paths pass has_prefill=False, preserving their existing shape-only behavior while exposing a single token-count API. This addresses the follow-up review on #51865: https://github.com/vllm-project/vllm/pull/51865#discussion_r3763278414 This does not duplicate an …
- fd04ede #51838 — [Refactor] Delete dead code in models (#51838)
- 作者: Wentao Ye | +34/-223 | 25 个文件
- is_internal_router is never used - some unused status for expert rank/range
🦀 Rust Frontend
- 3ee2df3 #51463 — [Frontend] Make
modeloptional on all/derenderrequest classes (#51463)- 作者: Vinay R Damodaran | +37/-14 | 3 个文件
Makes model optional on all four /derender request classes — DerenderChatRequest, DerenderCompletionRequest, DerenderChatStreamRequest, DerenderCompletionStreamRequest — and resolves the served model name server-side via request.model or self.models.model_name() when the field is omitted. Every comparable endpoint already treats model as optional: | Endpoint | model | | — | — | | /tokenize, /d…
⚡ Performance
- 7f9173d #50654 — [ROCm][Perf] Kimi-K3 Fused kernel for KDA decode (#50654)
- 作者: kliuae | +1592/-9 | 6 个文件
Kimi’s KDA layers do three things per decode step for every sequence and every value head: 1. Causal conv1d update over the Q/K/V projections, advancing a small per-sequence convolution state 2. Gated delta-rule recurrence, which updates a [head_dim, head_dim] recurrent state and produces the attention output 3. gated RMSNorm on that output On ROCm, currently these run as separate Triton kernels. …
- f97e502 #51738 — [Perf] Avoid more GPU<->CPU syncs on the model execution path (#51738)
- 作者: Nick Hill | +116/-56 | 15 个文件
Follow-on to the earlier sync-avoidance work, splitting further fixes out of the VLLM_GPU_SYNC_CHECK branch. Each of these removes a host roundtrip rather than suppressing it: - kv-sharing fast prefill: derive the fast-prefill decode metadata from host-side data. total_num_decode_tokens is just the number of logits indices, and the max per-request logits count is plumbed down from the runner (whic…
- 47ececb #48223 — [Perf][ROCm] Dual-stream decode with hipgraphs (#48223)
- 作者: Simon Danielsson | +55/-55 | 2 个文件
Fixes #48111. Enables (1) dual-stream decode for CUDA-like platforms with proper overlap (2) make them hip/cudagraph compatible. Only enabled on ROCm when using DP, as we observed performance regression under TP. Mutually exclusive with VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS. Disable with VLLM_DISABLE_SHARED_EXPERTS_STREAM=1 as usual. Gain: About -3-4% TPOT on 1k/1k and 8k/1k when using DPA…
- 5426311 #47896 — [Kernel][ROCm][Perf] FlyDSL decode-attention kernel for 4-bit TurboQuant KV cache (#47896)
- 作者: aditi-amd | +6396/-53 | 13 个文件
This PR upstreams the FlyDSL TurboQuant 4-bit KV-cache decode kernel developed and optimized by the AMD team. For background on the TurboQuant algorithm and its role in agentic vLLM serving, see our blog write-up [TurboQuant blog] (https://rocm.blogs.amd.com/artificial-intelligence/turboquant-vllm-agentic/README.html) This PR adds a custom FlyDSL TurboQuant 4-bit KV-cache decode kernel for ROCm / …
- 3fb7bb4 #51774 — [Perf] Avoid repeated multimodal prompt update scans (#51774)
- 作者: Tianyu Guo | +226/-121 | 2 个文件
Avoid quadratic prompt-update planning when many multimodal items share the same target. For non-empty replacements, the previous implementation applies one item per round. Each round still walks all unresolved items and rebuilds the match list, resulting in roughly O(N²) item processing. This change compiles updates with identical ordered (mode, target) choices into FIFO queues. Matching only exa…
🔩 Misc
- 5b5eae2 #49444 — [Misc] Enable test_silu_mul_fp8_quant_deep_gemm on XPU (#49444)
- 作者: pmanczak | +19/-17 | 2 个文件
persistent_masked_m_silu_mul_quant() called current_platform.get_device_capability().to_int() and asserted it wasn’t None. Device capability is a CUDA/ROCm concept, and XpuPlatform returns None, so the wrapper blew up on XPU instead of falling through to the Triton path. 1. batched_deep_gemm_moe.py - gate the C++ kernel on current_platform.is_cuda() and current_platform.has_device_capability(8…
✨ New Feature
- 6c95a64 #49315 — [2/N][Feat][Perf] Add new warmup infrastructure for JITs. Add predicate filtering for JIT warmup, and migrate Inkling FA4 (#49315)
- 作者: Roberto L. Castro | +743/-390 | 12 个文件
Description This PR migrates Inkling FA4 attention warmup to the shared JIT warmup contract introduced in #47451,
and deprecates the legacy CuTeDSL warmup path(there is a conflict with the Kimi K3 integration #50089 that prevents this deprecation from being completed. It will be addressed in a future PR). See https://github.com/vllm-project/vllm/issues/49349 for more context ### Motivation …