共 67 个 commit,涉及 407 个文件,+17198/-2867 行变动。
概要
| 统计项 | 数值 |
|---|---|
| Commit 数 | 67 |
| 变更文件 | 407 |
| 新增行数 | +17198 |
| 删除行数 | -2867 |
Commit 列表
🔩 Misc
- cdc4824 #48684 — [Misc] Remove
override_attention_dtype(#48684)- 作者: wangxiyuan | +0/-15 | 2 个文件
override_attention_dtype is only used for V0 and has been removd from https://github.com/vllm-project/vllm/pull/25351/ long time ago. It’s safe to remove it now. ## Test Result —
- 015660d #52145 — [Misc] Add missing return type annotations in outputs.py (#52145)
- 作者: Vineeta Tiwari | +13/-7 | 1 个文件
Purpose Add missing return type annotations to from_base() static methods and PoolingRequestOutput.repr() in vllm/outputs.py. All from_base() static methods on EmbeddingOutput, ClassificationOutput, ScoringOutput, EmbeddingRequestOutput, ClassificationRequestOutput, and ScoringRequestOutput lacked return type annotations. PoolingRequestOutput.repr() was missing -> str, making it the on…
📦 Other
- 1f7427b #52265 — [UT][XPU] fix b12x UT (#52265)
- 作者: Qiming Zhang | +24/-3 | 2 个文件
Intel CI failure: https://buildkite.com/vllm/intel-ci/builds/8908/canvas?sid=019ffe6c-17b4-443f-bc9f-8bf51457f9e2&tab=output Root cause: b12x.py does from vllm.utils.flashinfer import flashinfer_mxfp4_quantize, but the function was never registered in that module. The test’s monkeypatch.setattr then fails with AttributeError because there’s nothing to patch. Fix: Added flashinfer_mxfp4…
- 57bd0ed #51704 — [5/N][KV-Cache Layout Refactor] Backend-published KV packing via customize_spec (#51704)
- 作者: Lucas Wilkinson | +204/-219 | 18 个文件
Part of the KV-cache layout standardization series (RFC #42082). Stacked on #51612 — the diff shown includes it until that lands and this retargets main. Attention specs today carry quant-format sizing knowledge inline: nvfp4 / per-token-head branches in the page-size properties, a TQFullAttentionSpec subclass, and fp8_ds_mla constants in MLA spec overrides. This PR makes specs plain data and …
- 63a9a50 #52164 — [Attention][DSA] Take the native decode path for MTP=3 on SM90 (#52164)
- 作者: Zhuobin Huang | +232/-107 | 4 个文件
[Attention][DSA] Take the native decode path for MTP=3 on SM90 Closes #35878. The DSA indexer flattens a spec-decode batch into one single-token row per query whenever next_n falls outside {1, 2}, so with MTP=3 (next_n = 4) each request’s KV tile is read four times instead of once. DeepGEMM’s nv_dev branch, which vLLM already pins (cmake/external_projects/deepgemm.cmake), implements next_n = 4 o…
- 66728fe #49852 — [MRV2][Multimodal] Enable encoder cuda graph for model runner v2 (#49852)
- 作者: Isotr0py | +101/-26 | 4 个文件
- Enable encoder cuda graph on model runner v2. ## Test Result All tests should pass —
- 624999a #52138 — [XPU]bump up vllm_xpu_kernels to 0.1.13.2 (#52138)
- 作者: Kunshang Ji | +1/-1 | 1 个文件
Test Result —
- 3c8676a #51650 — [PP][XPU]Overlap async-scheduling PP sampled-token broadcast with compute (#51650)
- 作者: YiSheng5 | +11/-1 | 1 个文件
When async scheduling is enabled with pipeline parallelism (pp > 1), the sampled token ids produced by the last PP stage are sent back to the first stage via torch.distributed.broadcast on the PP device group. This broadcast was issued as a blocking collective (async_op=False) on the default compute stream. This PR makes that broadcast non-blocking (async_op=True) and defers the .wait() to _prepar…
- bda4c3e #51583 — [CPU] Fold the MXFP4 block scale in 2 instructions instead of 4 (#51583)
- 作者: ccaadaro | +135/-11 | 2 个文件
The AVX-512 MXFP4 unpack in csrc/cpu/sgl-kernels/vec.h applies the E8M0 block scale as an integer add on the bf16 exponent field. It has to keep the two zero codes at zero, because they have no exponent to shift, and that special case is written as and + cmpeq + add + blend — four instructions per vector. vptestmw sets a lane’s mask bit for exactly the lanes where (x & 0x7FFF) != 0, which is the c…
- 8e6d8e4 #52108 — [XPU][CI/Release][3/N] Add xpu wheel release to release pipeline (#52108)
- 作者: Kunshang Ji | +99/-29 | 6 个文件
Summary This PR adds XPU wheel building and publishing support to the Buildkite release pipeline, enabling pre-built XPU wheels to be distributed via wheels.vllm.ai. # Changes Release pipeline (release-pipeline.yaml) - Add a new Build wheel - x86_64 - XPU step that builds the XPU wheel using Dockerfile.xpu and uploads it to S3. - Rename the existing Publish XPU Triton shim index step to Publish …
- 6adad08 #51655 — Add Muse Glimmer model support (#51655)
- 作者: Tiezhen WANG | +4333/-16 | 21 个文件
Dense 29.6B vision-language model with a ViT-G/14 perception encoder and 128K context. Adds the model, its config and processor, channel-scoped reasoning and ATEM tool-call parsers, and DFlash speculative decoding support for its draft head. The model does not emit JSON tool calls and does not wrap reasoning in
tags. Every turn is a sequence of channel-scoped messages, and both parsers key… - 59d1af5 #49365 — Detect ROCm wheel variant from environment for precompiled wheels. (#49365)
- 作者: Aarushi Jain | +302/-12 | 2 个文件
Fixes AMD: Python-only Installation failing because ROCm precompiled wheels on wheels.vllm.ai use a different path layout than CUDA. - setup.py: Detect installed ROCm at runtime, match against published variants on wheels.vllm.ai/rocm/{commit}/, fall back to AMD PyPI. - python_only_compile.sh: Same variant resolution for the preflight metadata check.
- d18bb7b #52237 — [UT] fix device of test_outputs.py (#52237)
- 作者: Qiming Zhang | +11/-6 | 1 个文件
Use current_platform.device_type instead of device=“cuda” to enable the unit test to run successfully on all platforms.
- f80b66f #50062 — [Model Runner V2][Spec Decode] Add KV cache support for multi-layer MTP (#50062)
- 作者: Giancarlo Delfin | +196/-32 | 10 个文件
This PR adds the scheduler and KV-cache-manager support required for multi-module MTP (one MTP module per speculative step, e.g. Inkling’s 8-depth checkpoint). It is the companion to #48892, which introduced the speculator itself to Model Runner V2. The core property this PR protects: the multi-module drafter reads ahead of the computed tokens during prefill. MTP module m computes position p’s…
- b652ded #52148 — [Attention] Fix FlashInfer SM12x prefill with sinks (#52148)
- 作者: Andrii Skliar | +66/-24 | 2 个文件
Use FlashInfer’s sink-aware paged prefill wrapper on SM12x when attention sinks are enabled. The generic FA2 prefill path accepts a sinks argument but does not apply it, so #49718 can use XQA for decode while producing incorrect prefill output. The wrapper is specialized with the active dtypes, head dimensions, sliding window, and softmax scale. DCP, NVFP4, SM90/SM100, and sink-free paths are unch…
- 2e2ffd1 #52210 — [CI Failure] Fix CUDA wheel build for the Kimi K3 fused MLA kernel (#52210)
- 作者: Misha Goin | +31/-37 | 1 个文件
The release pipeline for building wheels has been failing since https://github.com/vllm-project/vllm/pull/51772 landed due to SM75 not supporting bf16 Latest failing job on main https://buildkite.com/vllm/release-v2/builds/5155/list?sid=019ffc97-c2e1-4216-87f4-9c55baf6e6fb&tab=output Use _typeConvert<scalar_t>::exists to discard unsupported packed conversion paths during host and pre-Ampere compil…
- 73b8394 #51633 — [Platform] Add check_runner_kv_caches_multi_layer (#51633)
- 作者: wangxiyuan | +34/-11 | 6 个文件
Add check_runner_kv_caches_multi_layer interface to avoid platform hardcode in bind_kv_cache. So that oot platform can override it to avoid error raising. ## Test Result —
- c5b7c06 #51159 — [ROCm] Defer
tilelangimport through its importfrom vllm.tilelang_utils import tilelangand relaxedhas_tilelang(#51159)- 作者: fxmarty-amd | +313/-52 | 8 个文件
Fixes https://github.com/vllm-project/vllm/issues/51151 This PR introduces vllm.tilelang_utils and prevents direct tilelang imports. Avoid importing TileLang during ROCm module import, because importing TileLang can load wrongful/bugged TVM and HIP stub symbols into the global process scope before AITER loads its JIT modules. This changes _tilelang_jit so ROCm applies tilelang.jit lazily on fi…
- 48825ac #51793 — [Quantization] Remove dead
QuantizationConfig.is_mxfp4_quant(#51793)- 作者: fxmarty-amd | +0/-41 | 4 个文件
This is dead code following https://github.com/vllm-project/vllm/pull/37128. This was originally added in https://github.com/vllm-project/vllm/pull/29008 that supported padding for gpt-oss / certain MXFP4 backends, see: https://github.com/xuebwang-amd/vllm/blob/c62f664e97977ee54ab1d1c77604ebb45081bc06/vllm/model_executor/layers/fused_moe/layer.py#L259-L277 This is now handled in: N/A ## Test Resul…
- 69d4c3a #52091 — Auto-ping Cohere on related issues (#52091)
- 作者: Cyrus Leung | +21/-0 | 2 个文件
As discussed offline ## Test Result —
- b245d8e #52173 — Apply logit softcapping in Transformers modelling backend (#52173)
- 作者: Harry Mellor | +3/-2 | 1 个文件
This is mainly used by Gemma models and only affects workloads where the exact logit values are important. Generation is unaffected because it does not reorder anything.
- 6014f9e #52079 — [Kimi-K3] Add GEMM-RS for sequence parallelism (#52079)
- 作者: Thien Tran | +1591/-13 | 11 个文件
Add GEMM-RS kernel for Blackwell, based on https://github.com/NVIDIA/cutlass/blob/dcf215a/examples/python/CuTeDSL/cute/blackwell/kernel/distributed/distributed_gemm_reduce_scatter_blackwell.py (multimem.ld_reduce) - Supports any value of M e.g. M=1023. However, only uses GEMM-RS when M>=128 since the kernel was not optimized for small/medium M. Only supports TP<=16, and requires all rank on the sa…
- 80d6d55 #52147 — Standardise weight tying on
ParallelLMHead.tie_weights(#52147)- 作者: Harry Mellor | +191/-214 | 56 个文件
vLLM expresses tied word embeddings in three different ways. Only one of them, self.lm_head = self.lm_head.tie_weights(embed_tokens), dispatches through quant_method.tie_weights. This PR converts the other two so tying is expressed one way everywhere. - self.lm_head.weight = embed_tokens.weight (33 sites) bypassed the quant method entirely, so it was wrong for quant methods that repack. The Parall…
- f3c1638 #51653 — [ROCm] Enable V2 model runner for Kimi-K3 on ROCm (#51653)
- 作者: vllmellm | +0/-16 | 1 个文件
Kimi-K3 on ROCm was gated from using V2 model runner. After validation using the up-to-date upstream. V2 model runner is working as expected. Command to start Kimi-K3 on mi355x ## Test Result Tested with lm-eval: | runner | run 1 | run 2 | |——–|——-|——-| | V1 | 96.66% | 97.04% | | V2 | 97.57% | 96.89% | —
- 95c9144 #42662 — [LoRA][Gemma4] Support vision tower LoRA (#42662)
- 作者: linitra24 | +309/-48 | 10 个文件
This PR adds the remaining LoRA plumbing needed for Gemma4 multimodal LoRA support. After #43798, Gemma4-MM vision linear layers are already converted through the Transformers backend path, so this PR no longer reimplements the Gemma4 vision tower. Instead, it focuses on the runtime LoRA mapping and token-counting pieces needed by Gemma4 image/video/audio inputs. Main changes: - Add a multimodal L…
- 8e1131e #52114 — [Model] [Quantization] Add Ling hybrid MXFP4 routed experts support (#52114)
- 作者: zexplorerhj | +48/-4 | 2 个文件
Add support for Ling checkpoints that use hybrid quantization: block FP8 for dense and shared-expert projections, and MXFP4 for routed experts. This change reads Ling-specific quantization metadata and remaps routed-expert scale names to the convention expected by Mxfp4MoEMethod for both the main and MTP models
- c4e9692 #50534 — [XPU] Add tuned Mamba SSU configs for Intel Arc Pro B70 (#50534)
- 作者: pmanczak | +247/-12 | 5 个文件
Add tuned selective_state_update configs for Intel Arc Pro B70 Graphics, and fix benchmarks/kernels/benchmark_selective_state_update.py so it runs on an XPU-only build. The config directory holds AMD and NVIDIA devices only, so Mamba/hybrid models on Intel GPUs fall back to the heuristic. Four files, two shapes: | shape | cache_dtype | exercised by | | — | — | — | | headdim=128,dstate=25…
- 5fee0a8 #51998 — chore: Upstream Cohere parser fixes + tests (#51998)
- 作者: JasonCohere | +533/-54 | 10 个文件
Adding some local fixes for Cohere parsers alongside corresponding tests
- 37c3bdf #50221 — fix(security): enforce audio decode duration limit in NanoNemotronVL (#50221)
- 作者: Juan Pérez de Algaba | +63/-2 | 2 个文件
The _extract_audio_from_videos method called load_audio_pyav without max_duration_s, allowing a small compressed video to decompress into gigabytes of PCM and crash the server via OOM. Pass VLLM_MAX_AUDIO_DECODE_DURATION_S to match the safeguard already used by AudioMediaIO. This should be merged only when https://github.com/vllm-project/vllm/pull/49948 branch) is merged
- b8baa31 #49458 — Hardware-agnostic model definition via HF transformer backend (1/N) (#49458)
- 作者: Thomas Ortner | +691/-5 | 10 个文件
This PR is an alternative approach to realize hardware-agnostic model definitions based on the HF transformer backend. In particular, the idea is that the modeling code of tail models resides in HF transformers and can be executed in vLLM through the help of the transformers backend, i.e., –model-impl transformers. The way this backend currently works is by replacing / patching particular layers …
- 903da60 #52134 — [Docs] Fix
WhisperEncoderLayer.forwarddocstring indots3_note(#52134)- 作者: Harry Mellor | +18/-9 | 1 个文件
The docs build emits griffe warnings for vllm/models/dots3_note/nvidia/audio_encoder.py: The WhisperEncoderLayer.forward docstring was inherited from the upstream HF Whisper implementation and never updated for this layer’s signature, which takes packed variable-length inputs (cu_seqlens_, max_seqlen_) and rotary embeddings instead of attention_mask/layer_head_mask. This documents the parameters…
- 1c3633a #52127 — [CI/Build][CPU] Shrink triton-cpu-build layer by dropping build artifacts (#52127)
- 作者: Li, Jiang | +3/-0 | 1 个文件
- The vllm-triton-cpu-build stage in docker/Dockerfile.cpu clones and builds triton-lang/triton-cpu, but only the resulting wheel is needed by later stages/the final image. - The cloned triton-cpu source/build tree and Triton’s downloaded LLVM/MLIR toolchain (under /root/.triton) were being retained in the final layer, growing the stage to 3.46GB. - This PR removes the triton-cpu source tree after…
- 10bcad2 #52123 — Update CODEOWNERS (#52123)
- 作者: Cyrus Leung | +2/-0 | 1 个文件
As discussed offline with @Isotr0py ## Test Result —
- 7bbbf7c #51251 — [Core] Configure custom encoder cache managers from VllmConfig (#51251)
- 作者: hotTea | +36/-8 | 4 个文件
Expose custom encoder cache manager configuration through VllmConfig for both online and offline inference. Custom encoder cache managers may require policy-specific parameters in addition to the encoder cache size. This PR provides a generic configuration path while preserving compatibility with existing built-in and constructor-only cache managers. This PR: - Adds an opaque manager_config field …
🐛 Bug Fix
- aa31003 #51664 — [Bugfix][Helm] Fix chart resource references (#51664)
- 作者: iwannagotobed | +84/-17 | 9 个文件
Assisted-by: Codex Fix inconsistent Helm chart resource references when custom labels and autoscaling are enabled. Before this change: - The Service selector used configured labels, while the Deployment selector and Pod labels were hard-coded to test/test. - The HPA targeted a non-existent Deployment named vllm. After this change, the Deployment selector, Pod labels, Service selector, and HPA targ…
- 20405bf #51989 — [Bugfix] Fix Cosmos3-Edge processor after transformers 5.15 release (#51989)
- 作者: bastefaniak | +20/-5 | 2 个文件
This PR fixes Cosmos3-Edge processor which is broken when transformers==5.15 is used, due to refactoring of underlying Qwen3-VL processor. With the fixes preprocessor will work correctly for both transformers==5.14 and 5.15. Also as model was released removed is_available_online=False from registry. ## Test Result —
- ac7509e #51099 — [Bugfix][CPU][RISC-V] Fix build: make FP32Vec copy constructors non-explicit (#51099)
- 作者: velonica0 | +6/-2 | 1 个文件
The RISC-V CPU backend has not compiled since #49021. cpu_attn_vec.hpp copy-initialises the halves of the pair returned by load_b_pair_vec. Copy-initialisation only considers non-explicit constructors, and declaring a copy constructor explicit also suppresses the implicitly-declared one. RISC-V declares both FP32Vec8 and FP32Vec16 copy constructors explicit(#36578), so overload resolution finds no…
- 653cc6f #50620 — [Bugfix][NIXL] Include transfer mode (push/pull) in the compatibility hash (#50620)
- 作者: Tzu-Ling Kan | +76/-3 | 8 个文件
Overview: Include the NIXL transfer mode (push vs pull) in the connector so a push (WRITE) connector and a pull (READ) connector can never be paired, and so an external router can distinguish them. Follow-up to #49230 (now merged), addressing review feedback from @iyastreb (#49230 thread, this PR’s thread). #### Details: The push (NixlPushConnector, WRITE) and pull (NixlConnector, READ) conne…
- c05d75a #50685 — [Bugfix][Refactor] Keep Qwen3Next layer boundaries sequence parallel (#50685)
- 作者: kzwrime | +75/-111 | 4 个文件
Fixes #50681. Qwen3.6-35B-A3B produces corrupted output during single-token decode when expert parallelism and MoE sequence parallelism are enabled with DP2 and TP2. The Qwen3Next model path infers the hidden-state layout from the first tensor dimension, but during TP2 single-token decode the full input and each padded local shard can all have one row. Shape alone therefore cannot determine whethe…
- fe4c5dc #52118 — [XPU] [Bugfix] process ragged weights in xpu linear backend (#52118)
- 作者: zofia | +29/-1 | 1 个文件
python examples/basic/offline_inference/generate.py –model gaunernst/DeepSeek-V2-Lite-Chat-FP8 –enforce-eager –max-model-len 2048 –trust-remote-code Before: After:
- b216db3 #51795 — [Bugfix] Reject negative token ids as out-of-vocabulary (#51795)
- 作者: Junhao Shen | +21/-0 | 2 个文件
InputProcessor._validate_model_input validates that caller-supplied token ids are within the vocabulary, but it only checks the upper bound, so a negative token id passes validation. A token id is used as an index downstream, and a negative index is never valid input. This PR adds a symmetric lower-bound check beside the existing upper-bound check. A negative id is now rejected with the same “out …
- 38f097f #51796 — [Bugfix] Reject NUL byte in structured_outputs.regex (#51796)
- 作者: Junhao Shen | +50/-0 | 3 个文件
A NUL byte is never meaningful in a regex pattern and is not handled by the regex-to-grammar conversion. Currently a structured_outputs.regex containing a NUL is passed through to the backend instead of being rejected. This PR rejects a regex containing a NUL at request validation, before backend selection, in SamplingParams._validate_structured_outputs — a clean HTTP 400 in every backend mode, in…
- 71b0da7 #52005 — [Bugfix] Fix …/mrope.py::apply_interleaved_rope() when torch.compile is used in torch==2.13 (#52005)
- 作者: bastefaniak | +52/-5 | 2 个文件
This PR fixes incorrect outputs of torch.compile …/mrope.py::apply_interleaved_rope() when it’s used with torch==2.13 (which newest vLLM uses), In torch==2.11 it worked correctly. We fix it by computing the same output in a way that torch.compile doesn’t break . This method does only indexing and assigning so compiling it should not introduce errors. Added test case, comparing eager to torch.com…
- 51def78 #52223 — [Bugfix] Reapply 50869 (#52223)
- 作者: Benjamin Chislett | +0/-24 | 1 个文件
#49969 accidentally reverted #50869 when resolving a merge conflict.
- 6355051 #45423 — [Bugfix] Correct prompt lengths for timed_traces benchmark (#45423)
- 作者: Stan Wozniak | +8/-4 | 3 个文件
#39795 introduced timed_traces support for vllm bench serve. The traces look as follows: Current implementation creates prompts with these input_lengths, and then runs: vLLM engine on the server side runs: Unfortunately, tokenizer logic is not idempotent, so whereas the client generates requests of length 6758, 7322, 7236, 2290, etc., the server receives requests of different length 7253, 7844, 76…
- 11c3fa4 #52139 — [Bugfix][ROCm][CI] Give the AITER MLA decode metadata stub its MLA dims (#52139)
- 作者: stefankoncarevic | +13/-3 | 1 个文件
Purpose tests/kernels/attention/test_rocm_aiter_mla_decode_metadata.py::test_persistent_decode_metadata_matches_fp8_golden fails on main with AttributeError: ’types.SimpleNamespace’ object has no attribute ‘q_lora_rank’. Two jobs report it, the dedicated AITER MLA job and the sharded kernels/attention job, but it is the same test. “[Model] Add native Dots3 NOTE multimodal support” (#51255), chan…
- e6b2a8a #50595 — [Bugfix][Structured Output] Mask request stop tokens in xgrammar until grammar terminates (#50595)
- 作者: yzong-rh | +41/-14 | 2 个文件
Following #49227’s merge, remove the patch used in HarmonyParser and update test to mirror production path. pytest tests/parser/test_harmony.py ## Test Result 64 passed, 38 warnings in 26.74s —
- 83d4c61 #52171 — [Bugfix] Declare SupportsEagle3 on KimiLinearForCausalLM (#52171)
- 作者: Nick Iusiumbeli | +10/-1 | 2 个文件
Purpose KimiK3ForConditionalGeneration (multimodal) declares SupportsEagle3; the text-only KimiLinearForCausalLM does not — even though both serve the same inner KimiLinearModel, which already inherits EagleModelMixin and implements the aux-hidden-state tap machinery. Serving a text-only Kimi-K3 checkpoint with EAGLE3-family speculative decoding (e.g. dspark) therefore dies at startup: Adding …
- 96acd47 #52122 — [Bugfix][MiniCPM-V] Fix AssertionError in get_dummy_mm_data when passing VideoDummyOptions to _get_dummy_images (#52122)
- 作者: Qiming Zhang | +18/-2 | 1 个文件
Issue: pytest tests/lora/test_minicpmv_tp.py::test_minicpmv_lora raises AssertionError on non-CUDA platforms (e.g., XPU). Root Cause: Commit 9a276d6375 added a runtime assertion to _get_dummy_images in dummy_inputs.py: assert overrides is None or isinstance(overrides, ImageDummyOptions) However, MiniCPMVDummyInputsBuilder.get_dummy_mm_data in minicpmv.py had always been passing video_overr…
- 64ca614 #47692 — [Bugfix] Fix
--data-parallel-start-rank 0being treated as unset increate_engine_config(#47692)- 作者: Ali Jaseem | +21/-2 | 2 个文件
EngineArgs.create_engine_config uses Python truthiness (if self.data_parallel_start_rank) instead of is not None to detect whether –data-parallel-start-rank was explicitly set. Since 0 is a valid, meaningful starting rank (the node owning the first slice of global DP ranks), an explicit –data-parallel-start-rank 0 is silently treated identically to “not specified.” This causes data_parallel_hybr…
- 2d24355 #52030 — [Bugfix] Fix packed GDN decode launch for large batch-head grids (#52030)
- 作者: Michael Goin | +33/-3 | 2 个文件
Avoid a CUDA launch failure in packed GDN decode when batch_size * num_value_heads exceeds the maximum CUDA grid Y/Z dimension of 65,535. The existing launch is preserved for normal sizes. Only overflowing cases use a split (value_tiles, value_heads, batch) grid. ## Test Result - Verified the failing Qwen shape (B=1024, HV=64, K=V=128) launches successfully. - Running vllm serve mgoin/Qwen3.8-2.4T…
- 170592a #52172 — [Bugfix] Disable sequence parallelism for Dots3 NOTE (#52172)
- 作者: 范裕达 | +3/-5 | 1 个文件
Why The DeepSeek V3.2 sequence-parallel refactor changed the inherited forward paths to use use_sequence_parallel. Dots3 NOTE uses custom model and decoder initializers and has not adopted the new sequence-parallel execution path. This causes serving to fail during KV cache profiling with: AttributeError: ‘Dots3NoteModel’ object has no attribute ‘use_sequence_parallel’ ## What changed Explicitl…
- d0ae25e #52021 — [Bugfix] Preserve Anthropic disable_parallel_tool_use (#52021)
- 作者: Taneem Ibrahim | +4/-0 | 2 个文件
Anthropic’s disable_parallel_tool_use was silently discarded, leaving the converted OpenAI request with parallel_tool_calls=True. This preserves the field and maps it to the existing inverse OpenAI setting. ## Reproducer On Main On this branch ## Test Plan and Results
- 399f974 #50874 — [Bugfix][R3] Size monolithic routing replay buffer for DP (#50874)
- 作者: TomerBN-Nvidia | +30/-5 | 2 个文件
Fix routing-replay capture for the FlashInfer monolithic MoE kernel under naive data parallelism, including padded sequence-parallel shards when expert parallelism is enabled. Two related assumptions fail in a TP2/DP2 deployment: 1. Replay buffer capacity. max_num_tokens is a per-rank scheduler limit, while the naive dispatch path all-gathers rank-local batches before invoking the monolithic k…
📖 Documentation
- 69e0e58 #52289 — [Doc] Update model support information (#52289)
- 作者: Jee Jee Li | +18/-5 | 2 个文件
Test Result —
- 443fa5a #51611 — [Doc] Fix stale rejection_sample_method and synthetic_acceptance_rate (#51611)
- 作者: QWERQWERQWE86 | +3/-2 | 1 个文件
Fixes #51609 Sync the –speculative-config table in docs/features/speculative_decoding/README.md with the current code: 1. rejection_sample_method (line 87): strict, probabilistic, synthetic (default strict) -> standard, synthetic, block (default standard); probabilistic now belongs to draft_sample_method. See #40651. 2. synthetic_acceptance_rate (line 88): split into synthetic_acceptance_rates (l…
🦀 Rust Frontend
- b8165e5 #52261 — [Frontend] Consolidate entrypoint exception handler (#52261)
- 作者: wang.yuqi | +418/-343 | 31 个文件
Consolidate entrypoint exception handler Part of #52131 (Move api_server.py out openai folder) pytest tests/entrypoints/serve/exception_handler/ ## Test Result pass —
- 7553aac #51906 — [Frontend] Add routed-experts prompt offset (#51906)
- 作者: aoshen02 | +114/-61 | 13 个文件
- Add routed_experts_prompt_start to OpenAI chat/completion requests and SamplingParams, allowing clients to omit an already-known prompt prefix from returned R3. - Centralize NumPy-to-base64 serialization used by existing R3 responses and document the int32 expert-ID representation. - Keep OpenAI streaming behavior unchanged: R3 remains supported only on existing non-streaming responses. ## Why t…
- 152c913 #52098 — [Frontend] Log output token IDs at DEBUG level (#52098)
- 作者: yang rui | +52/-25 | 3 个文件
Allow operators to keep human-readable generated output logs without emitting output token IDs at the default INFO level. Following maintainer feedback, this now mirrors the existing request-input logging split instead of adding a new CLI flag: - INFO keeps generated text and the finish reason. - DEBUG additionally logs output token IDs. - –max-log-len continues to truncate both output text and t…
⚡ Performance
- 103c419 #52277 — [Perf][Frontend] Vectorize Cohere binary embedding bit-packing (#52277)
- 作者: Fangchen Li | +57/-24 | 2 个文件
_pack_binary_embeddings (vllm/entrypoints/pooling/embed/protocol.py) bit-packs embeddings for the Cohere /v2/embed binary / ubinary embedding types with a nested Python loop. This PR replaces it with np.packbits, which is ~4.3x faster with byte-identical output. ## Test Result All tests passed. ~4x performance improvemtn on M1 mac. —
- 1be3628 #51674 — [Kernel][Perf] Add fused CUDA post-conv MTP decode kernel for Qwen3.5 GDN (#51674)
- 作者: Jie Fang | +1456/-7 | 12 个文件
Speed up Qwen3.5 (GDN linear attention) MTP speculative decode on Blackwell. During MTP decode, the Triton path launches a chain of small kernels per step (gating, delta-rule recurrence, state rewind/update, gated RMSNorm), which leaves the GPU latency-bound at decode batch sizes. This PR adds a single fused CUDA kernel, fused_gdn_decode_post_conv_mtp, that consumes the post-convolution mixed_qkv …
- f962616 #51862 — [ROCm][Perf] Kimi-K3 Remove prefill pipeline stall in chunk KDA (#51862)
- 作者: kliuae | +188/-43 | 5 个文件
On ROCm’s Kimi-K3 path, each prefill/mixed step has a stall in prepare_chunk_indices, caused by two interacting factors: - tolist() D2H copy makes the host wait until the device queue is emptied - Host ops executing before the next H2D copy, while the device is sitting idle The data that the host waits on is already sitting on the host. At each step, GDNAttentionMetadataBuilder builds chunk_indice…
🧪 CI/Tests
- d4c24e6 #52252 — [CI] Increase extended generation test timeout (#52252)
- 作者: Lucas Wilkinson | +1/-1 | 1 个文件
Raise the Language Models Test (Extended Generation) timeout from 65 to 80 minutes. This keeps the existing test coverage and single-H200 resource shape unchanged. ## Why PR #48186 reduced this timeout from 110 to 65 minutes based on one successful 48.9-minute nightly and a 15-minute buffer. Runtime variance has since consumed that buffer: - main #78476 passed in 64m57s. - Main timed out on Ju…
- b96bcd0 #51280 — [ROCm][CI] Solidify entrypoint LLM lifecycle (#51280)
- 作者: Andreas Karatzas | +974/-1075 | 24 个文件
- Replace direct LLM(…) construction throughout tests/entrypoints with the shared VllmRunner lifecycle. - Add an ExitStack-backed runner factory for tests that need one long-lived runner or several concurrent runners. - Consolidate multimodal, structured-output, offline-mode, collective-RPC, pooling, and weight-transfer cleanup onto the complete runner shutdown path. - Defer ROCm VRAM settling w…
🖥️ Kernel
- 3c79b1a #52016 — [Kernel] Add B12X dense linear backends (#52016)
- 作者: Luke Alonso | +2753/-3 | 17 个文件
This PR integrates B12X dense linear kernels for NVIDIA SM120 and SM121 GPUs through the existing vLLM linear backend interfaces. B12X is an optional dependency installed with vllm[b12x] and pinned to b12x==1.2.4; it is a pure-Python CuTe DSL package and requires no additional vLLM build step. Supported linear paths are: - Per-tensor FP8. - 128x128 block-scaled FP8. - MXFP8. - NVFP4 and MXFP4. B12…
- 827a2af #48666 — [Kernel] Gemma-4 FA4 FP8 Kernel (#48666)
- 作者: Jhao-Ting Chen | +152/-29 | 8 个文件
Gemma-4 uses 256-wide heads in sliding_attention and 512-wide heads in full_attention. On SM90, the full-attention layers upgrade from FA3 to the FA4 CuTeDSL kernel. This PR wires the FA4 FP8-KV-dequant path from vllm-project/flash-attention#164 into vLLM, allowing Gemma-4 to use FP8 KV cache across both FA3 sliding attention and FA4 full attention. It also preserves FP8 KV-cache scales for Gemma-…
✨ New Feature
- 50ba4bc #49577 — [Feature] Mask Replay (#49577)
- 作者: vx120 | +556/-5 | 24 个文件
This PR adds experimental support for sampling distribution replay. A sampling mask represents the vocabulary support retained after top-k/top-p filtering. It is not an attention mask and does not affect causal attention or KV-cache behavior. When enabled, vLLM returns the sampling support for each generated token in a CSR-style representation: The feature is opt-in and does not change default…