共 38 个 commit,涉及 232 个文件,+4879/-1746 行变动。
概要
| 统计项 | 数值 |
|---|---|
| Commit 数 | 38 |
| 变更文件 | 232 |
| 新增行数 | +4879 |
| 删除行数 | -1746 |
Commit 列表
📦 Other
- d480199 #51901 — [CI/Build] Add warning for unsupported global PTX architecture requests in… (#51901)
- 作者: Shane Widanagama | +63/-5 | 4 个文件
… CMake configuration. Implements one item from #9129: Warn that PTX builds are not currently supported (post [CI/Build] Per file CUDA Archs (improve wheel size and dev build times) #8845), currently if there is a +PTX in TORCH_CUDA_ARCH_LIST this will be ignored. We should warn when this is the case Users can request PTX through TORCH_CUDA_ARCH_LIST values such as 8.0+PTX. vLLM strips the Torch…
- 4215646 #52384 — [Rust Frontend][gRPC] Preserve skip_special_tokens decoding option (#52384)
- 作者: Biswa Panda | +5/-1 | 2 个文件
Align the native Generate gRPC API with vLLM’s existing Python serving APIs by carrying the per-request skip_special_tokens decoding option through ResponseOptions. The Python chat completions, completions, responses, and token-native APIs default this option to true, but the gRPC request did not expose it. Consequently, a gRPC caller could not explicitly preserve tokenizer-defined special markers…
- 615d4cf #43107 — [Core] Check for GPU<->CPU syncs during CI (#43107)
- 作者: Nick Hill | +607/-364 | 44 个文件
vLLM now uses asynchronous scheduling by default and in the majority of cases. Performance relies on the absence of any gpu<->cpu synchronizations on the main cuda stream, but such syncs can be opaque and it is easy for them to creep in accidentally. This change adds a VLLM_GPU_SYNC_CHECK env var which enables torch.cuda.set_sync_debug_mode for the model forward pass and sampler, so that we can ea…
- d6f17f3 #52374 — [MRV2] Support attention-free models (#52374)
- 作者: Nick Hill | +7/-2 | 3 个文件
These use MambaHybridModelState And other minor changes for CI compatibility when switching to MRV2 by default. These are split out from https://github.com/vllm-project/vllm/pull/46646.
- 9df9b0b #52400 — [ROCm]: Drop pybind11 from Dockerfile.rocm to prevent version mismatch (#52400)
- 作者: Rohan Potdar | +1/-1 | 1 个文件
Currently, the AITER kernels in Dockerfile.rocm_base were built with pybind11 3.0.4, whereas the pybind11 install in Dockerfile.rocm installs pybind 3.1.0 (released 08/08/2026). The API incompatibility breaks Kimi K2.6/K3 among other models: https://github.com/ROCm/aiter/issues/4770 Quick fix: remove the pybind11 rebuild from the latter; we’ll eventually unify the Dockerfiles as part of the ROCk e…
- c794754 #50597 — [ROCm]Remove special-case SiTU support model-specific gating (#50597)
- 作者: sroberts-amd | +97/-146 | 3 个文件
Mxfp4MoEMethod previously contained a model-specific predicate (_use_k3_situ_aiter) that special-cased the Kimi-K3 SiTU activation, gating three separate code paths — backend selection in init, size round-up bypass in maybe_roundup_sizes, and an entirely separate weight-shuffle method _setup_kernel_k3_situ in process_weights_after_loading. This tied SiTU behavior to a specific model identity r…
- f473870 #51316 — [Rust Frontend][gRPC] Add RL lifecycle control (#51316)
- 作者: Connor Carpenter | +614/-16 | 17 个文件
Add reinforcement-learning lifecycle operations to the Rust frontend’s existing gRPC Control service. - Add pause and resume RPCs for scheduler control. - Add sleep and wake RPCs for GPU memory management. - Add weight-transfer initialization, start, update, finish, and version RPCs. - Advertise RL capabilities through the EngineCore ready handshake and reject unsupported operations before dispatc…
- 9b0ab5d #51216 — [ROCm][AMD] Enable preshuffled sparse indexing for 16-token blocks (#51216)
- 作者: James E T Smith | +16/-2 | 3 个文件
Allow ROCm sparse MLA indexer backends to use KV-cache block sizes aligned to 16 tokens. Previously, –block-size 16 selected kernel block size 1, disabling AITER’s preshuffled FP8 paged-MQA path. The preshuffled AITER kernel was more performant than the non-shuffled variant so we were leaving some performance on the table. This PR fixes that. AI assistance was used to prepare this change. Profile…
- 03a8d0b #50487 — [Model][Spec Decode] Tap the pre-norm AttnRes mixture as the Kimi K3 DFlash aux state (#50487)
- 作者: Rahul Chalamala | +350/-1 | 4 个文件
The DFlash drafter consumes auxiliary hidden states captured at a fixed set of target layers. K3 captures the post-mixture stream, which is not what the drafter was trained against: the AttnRes residual mixture is applied before the layer norm, and the current capture site reads the value after it. Tapping the pre-norm mixture instead recovers the stream the drafter expects. This selects the corre…
- 1f7427b #52265 — [UT][XPU] fix b12x UT (#52265)
- 作者: Qiming Zhang | +24/-3 | 2 个文件
Intel CI failure: https://buildkite.com/vllm/intel-ci/builds/8908/canvas?sid=019ffe6c-17b4-443f-bc9f-8bf51457f9e2&tab=output Root cause: b12x.py does from vllm.utils.flashinfer import flashinfer_mxfp4_quantize, but the function was never registered in that module. The test’s monkeypatch.setattr then fails with AttributeError because there’s nothing to patch. Fix: Added flashinfer_mxfp4…
- 57bd0ed #51704 — [5/N][KV-Cache Layout Refactor] Backend-published KV packing via customize_spec (#51704)
- 作者: Lucas Wilkinson | +204/-219 | 18 个文件
Part of the KV-cache layout standardization series (RFC #42082). Stacked on #51612 — the diff shown includes it until that lands and this retargets main. Attention specs today carry quant-format sizing knowledge inline: nvfp4 / per-token-head branches in the page-size properties, a TQFullAttentionSpec subclass, and fp8_ds_mla constants in MLA spec overrides. This PR makes specs plain data and …
- 63a9a50 #52164 — [Attention][DSA] Take the native decode path for MTP=3 on SM90 (#52164)
- 作者: Zhuobin Huang | +232/-107 | 4 个文件
[Attention][DSA] Take the native decode path for MTP=3 on SM90 Closes #35878. The DSA indexer flattens a spec-decode batch into one single-token row per query whenever next_n falls outside {1, 2}, so with MTP=3 (next_n = 4) each request’s KV tile is read four times instead of once. DeepGEMM’s nv_dev branch, which vLLM already pins (cmake/external_projects/deepgemm.cmake), implements next_n = 4 o…
- 66728fe #49852 — [MRV2][Multimodal] Enable encoder cuda graph for model runner v2 (#49852)
- 作者: Isotr0py | +101/-26 | 4 个文件
- Enable encoder cuda graph on model runner v2. ## Test Result All tests should pass —
- 624999a #52138 — [XPU]bump up vllm_xpu_kernels to 0.1.13.2 (#52138)
- 作者: Kunshang Ji | +1/-1 | 1 个文件
Test Result —
- 3c8676a #51650 — [PP][XPU]Overlap async-scheduling PP sampled-token broadcast with compute (#51650)
- 作者: YiSheng5 | +11/-1 | 1 个文件
When async scheduling is enabled with pipeline parallelism (pp > 1), the sampled token ids produced by the last PP stage are sent back to the first stage via torch.distributed.broadcast on the PP device group. This broadcast was issued as a blocking collective (async_op=False) on the default compute stream. This PR makes that broadcast non-blocking (async_op=True) and defers the .wait() to _prepar…
- bda4c3e #51583 — [CPU] Fold the MXFP4 block scale in 2 instructions instead of 4 (#51583)
- 作者: ccaadaro | +135/-11 | 2 个文件
The AVX-512 MXFP4 unpack in csrc/cpu/sgl-kernels/vec.h applies the E8M0 block scale as an integer add on the bf16 exponent field. It has to keep the two zero codes at zero, because they have no exponent to shift, and that special case is written as and + cmpeq + add + blend — four instructions per vector. vptestmw sets a lane’s mask bit for exactly the lanes where (x & 0x7FFF) != 0, which is the c…
🐛 Bug Fix
- 97388c4 #51538 — [Bugfix] Make DSV4 sparse MLA work end-to-end for plain decode, MTP, and DSpark (#51538)
- 作者: Gabriel Wu | +797/-120 | 20 个文件
DeepSeek-V4-Flash-0731 could not run reliably through the SM120 sparse MLA backend. This fixes the seven defects that blocked it across all three decode modes – plain decode, MTP, and DSpark – verified end-to-end on 8xRTX PRO 6000 Blackwell across in-flight batching and prefill/decode disaggregation. Commits 1-5 unblock DSpark. Commits 6-7 fix a hang that is not DSpark-specific: it strands a…
- 5cecfc0 #52431 — [Bugfix] Fix modelscope usage (#52431)
- 作者: Cyrus Leung | +3/-5 | 2 个文件
- Fix a KeyError when calling modelscope_list_repo_files, because recent versions of ModelScope no longer return the “Type” field. (I encountered this issue when trying to load meta-models/Muse-Glimmer-30B via ModelScope) - Accept VLLM_USE_MODELSCOPE=1, not just VLLM_USE_MODELSCOPE=True, to be consistent with other boolean-based env vars. ## Test Result —
- acb0f1d #52288 — [Bugfix][Spec Decode] DSpark: inherit the target’s attention backend when the speculative config names none (#52288)
- 作者: Yongye Zhu | +7/-1 | 1 个文件
load_dspark_model passes backend=speculative_config.attention_backend when building the draft’s config. That field is None unless the user names a backend inside –speculative-config, and None re-runs backend auto-selection rather than meaning “same as target” — so an explicitly pinned target backend is silently discarded and the draft can pick a different attention class. On DeepSeek V4 with –at…
- e078a22 #52241 — [Bugfix] Widen flashinfer.comm import guard so a failed import doesn’t abort engine startup (#52241)
- 作者: shanjiaz | +2/-2 | 1 个文件
The optional flashinfer.comm import in allreduce_rms_fusion.py is guarded by except ImportError, so any other exception during import aborts EngineCore startup. Hit with flashinfer-python==0.6.16.post3 on Python 3.11, which raises TypeError at import time — killing startup on a tp_size=1 run with fuse_allreduce_rms: False. Widen the guard to except Exception and log a warning, so a broken flashinf…
- aa31003 #51664 — [Bugfix][Helm] Fix chart resource references (#51664)
- 作者: iwannagotobed | +84/-17 | 9 个文件
Assisted-by: Codex Fix inconsistent Helm chart resource references when custom labels and autoscaling are enabled. Before this change: - The Service selector used configured labels, while the Deployment selector and Pod labels were hard-coded to test/test. - The HPA targeted a non-existent Deployment named vllm. After this change, the Deployment selector, Pod labels, Service selector, and HPA targ…
- 20405bf #51989 — [Bugfix] Fix Cosmos3-Edge processor after transformers 5.15 release (#51989)
- 作者: bastefaniak | +20/-5 | 2 个文件
This PR fixes Cosmos3-Edge processor which is broken when transformers==5.15 is used, due to refactoring of underlying Qwen3-VL processor. With the fixes preprocessor will work correctly for both transformers==5.14 and 5.15. Also as model was released removed is_available_online=False from registry. ## Test Result —
🦀 Rust Frontend
- ac2ae87 #45802 — [Frontend] Support count_reasoning_tokens in the Streaming Parser Engine (#45802)
- 作者: Chauncey | +669/-73 | 16 个文件
Add token-aware reasoning token counting for the Streaming Parser Engine and surface the count through OpenAI-compatible usage fields. - Adds completion_tokens_details.reasoning_tokens to usage responses. - Propagates token counts through the parser engine pipeline: TokenIDScanner -> IncrementalLexer -> StreamingParserEngine -> SemanticEvent. - Counts only REASONING_CHUNK tokens, excluding reasoni…
- b8165e5 #52261 — [Frontend] Consolidate entrypoint exception handler (#52261)
- 作者: wang.yuqi | +418/-343 | 31 个文件
Consolidate entrypoint exception handler Part of #52131 (Move api_server.py out openai folder) pytest tests/entrypoints/serve/exception_handler/ ## Test Result pass —
🧪 CI/Tests
- 44fc57d #49515 — [ROCm][CI] Select CPU platform for native no-GPU jobs (#49515)
- 作者: Andreas Karatzas | +60/-15 | 3 个文件
Motivation: Basic Models Test (Other CPU) and Async Engine, Inputs, Utils, Worker, Config (CPU) The native no_gpu: true jobs correctly hide MI300 devices, but the reused ROCm wheel suffix prevented platform probing from selecting CPU. A native job expecting zero GPUs now exports VLLM_TARGET_DEVICE=cpu, and explicit CPU selection takes precedence over wheel metadata and host accelerator probes. Thi…
- bb4b448 #52326 — [CI] Shard Humming H100 eval (#52326)
- 作者: Kevin H. Luu | +17/-2 | 4 个文件
Why lm-eval-humming-f16-h100 took 62.6 minutes in full CI build 83851. Its 14 model configurations ran serially even though each evaluation is independent. This does not duplicate an open pull request. Open PRs were searched by the exact step key and Humming/H100 sharding terms before implementation. ## What changed - run the step with three Buildkite jobs - partition all 14 configurations into…
- 81e81da #52327 — [CI] Shard MoE refactor B200 eval (#52327)
- 作者: Kevin H. Luu | +22/-2 | 5 个文件
Why moe-refactor-integration-test-b200-temporary took 81.3 minutes in full CI build 83851. Its 19 model configurations ran serially even though each evaluation is independent. This does not duplicate an open pull request. Open PRs were searched by the exact step key and MoE/B200 sharding terms before implementation. ## What changed - run the step with four Buildkite jobs - partition all 19 conf…
- 694db07 #52323 — [CI] Shard multimodal extended generation 2 (#52323)
- 作者: Kevin H. Luu | +3/-2 | 1 个文件
- run multi-modal-models-extended-generation-2 as four deterministic pytest shards - label the parallel jobs with their shard number - preserve the existing split(group=0) and not core_model selection ## Why Buildkite build #83851 took 71 minutes for this job. Its recent passed-run distribution is 68 minutes median, 75 minutes p90, and 131 minutes max. Four shards provide p90 headroom and reduce s…
- 549cef0 #52322 — [CI] Shard extended pooling model tests (#52322)
- 作者: Kevin H. Luu | +6/-2 | 1 个文件
- run language-models-test-extended-pooling as four deterministic pytest shards - label the parallel jobs with their shard number - keep the AMD mirror at one unsharded job with its original command ## Why Buildkite build #83851 took 72 minutes for this job. Its recent passed-run distribution is 70 minutes median, 82 minutes p90, and 112 minutes max. Four shards leave substantially more headroom t…
- d87ef45 #52328 — [CI] Shard Quantization job into 4 parallel shards (≤30 min target) (#52328)
- 作者: Kevin H. Luu | +4/-3 | 1 个文件
What Shard the long-pole Quantization CI job (single job, ~71–75 min wall in build #83851; 90-day 68/73/75 min) into 4 parallel Buildkite shards so each shard lands ≤30 min. Rebased onto current main, preserving the live device/key/V2 env/torchao 0.17+CUDA13 pin and the -k ’not test_compressed_tensors_w4a8_fp8’ selection. ## How - parallelism: 4 on the Quantization step + pytest-shard (–shard-…
- 925ea7e #50589 — [CI][Test] Seed the DeepEP v2 MoE workers, not just the parent (#50589)
- 作者: Guanxin Li | +5/-0 | 1 个文件
Fixes the nondeterminism behind #50184, where kernels/moe/test_deepep_v2_moe.py::test_deep_ep_v2_moe fails intermittently with: The test seeds the wrong process. test_deep_ep_v2_moe calls set_random_seed(7) at line 328, which runs in the parent and covers make_test_weights(). But the tensors the assertion actually depends on are built in the worker — _deep_ep_v2_moe() calls TestTensors.make(co…
- 83ded8d #52331 — [Test][LoRA] Speed up the LoRA test job (#52331)
- 作者: stefankoncarevic | +104/-126 | 4 个文件
The LoRA job is one of the slower CI jobs and most of its time goes into work that is not under test. This PR removes three such costs, in test code only; no test is dropped or weakened. 1. Run the punica reference on device instead of CPU test_punica_ops.py PR #47534 replaced the torch.einsum reference with a per-LoRA matmul loop and put it on CPU, to avoid materializing a large lora_weight[e…
- d4c24e6 #52252 — [CI] Increase extended generation test timeout (#52252)
- 作者: Lucas Wilkinson | +1/-1 | 1 个文件
Raise the Language Models Test (Extended Generation) timeout from 65 to 80 minutes. This keeps the existing test coverage and single-H200 resource shape unchanged. ## Why PR #48186 reduced this timeout from 110 to 65 minutes based on one successful 48.9-minute nightly and a 15-minute buffer. Runtime variance has since consumed that buffer: - main #78476 passed in 64m57s. - Main timed out on Ju…
⚡ Performance
- 7b544ec #49793 — [Spec Decode][Perf] Fuse the MTP trailing all-reduce; local-argmax draft tokens (#49793)
- 作者: xiaozhoupy | +33/-5 | 1 个文件
Two optimizations on the DeepSeek-V3.2 / GLM-5.2 MTP draft path. - Fuse the trailing all-reduce into the final RMSNorm on the non-sequence-parallel path, as the main model already does at layer boundaries. The sequence-parallel path is unchanged. - Greedy draft tokens via vocab-parallel local argmax (get_top_tokens), skipping the full-vocab all-gather in compute_logits. The proposer alread…
- 3e3ceb1 #52369 — [Perf] Avoid more GPU<->CPU syncs in multimodal encoders (#52369)
- 作者: Nick Hill | +81/-72 | 9 个文件
Host data staged across blocking, now non-blocking or via async_tensor_h2d: - keye: position ids / cu_seqlens / sample indices on both the image and video paths, and the same pattern in paddleocr_vl, which carries a copy of that encoder. - ernie45_vl: the rotary pos_ids gather (a device tensor was being indexed with a CPU one), the encoder cu_seqlens, and the two numpy-built slice-index tensors in…
- 103c419 #52277 — [Perf][Frontend] Vectorize Cohere binary embedding bit-packing (#52277)
- 作者: Fangchen Li | +57/-24 | 2 个文件
_pack_binary_embeddings (vllm/entrypoints/pooling/embed/protocol.py) bit-packs embeddings for the Cohere /v2/embed binary / ubinary embedding types with a nested Python loop. This PR replaces it with np.packbits, which is ~4.3x faster with byte-identical output. ## Test Result All tests passed. ~4x performance improvemtn on M1 mac. —
🔩 Misc
- cdc4824 #48684 — [Misc] Remove
override_attention_dtype(#48684)- 作者: wangxiyuan | +0/-15 | 2 个文件
override_attention_dtype is only used for V0 and has been removd from https://github.com/vllm-project/vllm/pull/25351/ long time ago. It’s safe to remove it now. ## Test Result —