共 86 个 commit,涉及 360 个文件,+25362/-3348 行变动。

概要

统计项数值
Commit 数86
变更文件360
新增行数+25362
删除行数-3348

Commit 列表

⚡ Performance

  • 5426311 #47896 — [Kernel][ROCm][Perf] FlyDSL decode-attention kernel for 4-bit TurboQuant KV cache (#47896)
    • 作者: aditi-amd | +6396/-53 | 13 个文件

    This PR upstreams the FlyDSL TurboQuant 4-bit KV-cache decode kernel developed and optimized by the AMD team. For background on the TurboQuant algorithm and its role in agentic vLLM serving, see our blog write-up [TurboQuant blog] (https://rocm.blogs.amd.com/artificial-intelligence/turboquant-vllm-agentic/README.html) This PR adds a custom FlyDSL TurboQuant 4-bit KV-cache decode kernel for ROCm / …

  • 3fb7bb4 #51774 — [Perf] Avoid repeated multimodal prompt update scans (#51774)
    • 作者: Tianyu Guo | +226/-121 | 2 个文件

    Avoid quadratic prompt-update planning when many multimodal items share the same target. For non-empty replacements, the previous implementation applies one item per round. Each round still walks all unresolved items and rebuilds the match list, resulting in roughly O(N²) item processing. This change compiles updates with identical ordered (mode, target) choices into FIFO queues. Matching only exa…

  • 0914ed2 #51725 — [Perf] Adaptive budget for spec scheduled input tokens, ~60% better Kimi K3 DSpark TTFT (#51725)
    • 作者: Wentao Ye | +54/-24 | 4 个文件

    Let’s consider max_num_seqs=1024and max_num_batched_tokens=8192 (by default) Actual Request num | Old logic scheduled tokens | Now – | – | – 1 | 2048 | 8186 32 | 2048 | 8000 128 | 2048 | 7424 1024 | 2048 | 2048 This PR make the scheduled tokens much larger when request num is small by using an adaptive strategy This is extremely helpful when request num is small and give very huge perf impr…

  • 8977ea8 #50333 — [Perf] Skip detokenization in offline beam search (#50333)
    • 作者: Myeongwon Kim | +2/-0 | 1 个文件

    What & why Follow-up to #46422, which skipped detokenization in online beam search and noted the offline path has the identical pattern. Offline beam search asks for logprobs = 2 * beam_width on every internal per-step request. The engine detokenizes all of those token ids into strings on every step, but beam search never reads them. It ranks beams by cum_logprob and decodes the final text itse…

  • 405bc86 #51507 — [Perf] Launch the top-k/top-p Triton sampler kernel with 8 warps (#51507)
    • 作者: kiroxu | +7/-0 | 1 个文件

    _topk_topp_kernel (the Qrita sampler kernel from #42191) runs one program per logits row, and each program serially sweeps the whole vocab row in BLOCK_SIZE=8192 tiles. Triton’s default num_warps=4 leaves an 8192-wide fp32 tile at 16 elements per lane, so per-tile latency — which directly bounds kernel latency — is warp-starved. This kernel is on the hot path for seeded / per-request-generator sam…

  • 79c865b #51430 — [Perf] Narrow DeepSeek V4 eager CUDA graph region (#51430)
    • 作者: Woosuk Kwon | +90/-97 | 3 个文件

    Narrow the DeepSeek V4 eager CUDA graph region so attention input preparation remains captured, and organize the flow so the execution stages are visible in one linear forward method. The existing attention_impl eager break includes Q up-projection, fused Q normalization/RoPE/KV insertion, indexer input preparation, and MLA/indexer compression. This change moves those producer operations into forw…

  • fac808b #49436 — [Perf][Hybrid] 3D-grid tiling of the state-copy Triton kernels (#49436)
    • 作者: Francesco Fusco | +347/-68 | 3 个文件

    Follow-up to https://github.com/vllm-project/vllm/pull/48110. This PR further improves the performance by adding a 3D grid and lift the hard 8B alignment precondition by using a head/body/tail pattern. Those are the main changes: 1. Tiles the temporal copy across CTAs. The uint64 body range of each temporal state is partitioned across TEMPORAL_TILES CTAs along a new program_id(2) axis. Grid go…

🧪 CI/Tests

  • d648236 #51806 — [XPU][CI] Fix ExampleConnector KV cache device selection (#51806)
    • 作者: liuzhenwei | +3/-3 | 3 个文件
    • Load KV cache tensors onto the destination cache tensor’s device instead of hard-coding CUDA. - Add example connector test into XPU CI. ## Test Result —
  • 75903e5 #51666 — [CI][AMD] Persist the openai-harmony tiktoken vocab cache across jobs (#51666)
    • 作者: stefankoncarevic | +44/-5 | 2 个文件

    gpt-oss entrypoints tests can fail with a 500 in the middle of a run: Seen in amd-ci #11876 (mi355_1: Entrypoints Integration (API Server OpenAI - Part 2), commit 243c63baf), where 4 of the 8 TestGPTOSSChat cases failed. The server started fine; the very first POST /v1/chat/completions failed while rendering the harmony prompt. openai_harmony does not ship the tiktoken vocab. It downloads it on fi…

  • 1a17273 #50713 — [CI] Solidify speculative decoding E2E coverage (#50713)
    • 作者: Andreas Karatzas | +749/-471 | 18 个文件
    • Standardize the reorganized spec-decode tests on vllm_runner for consistent setup, defaults, and cleanup. - Add AMD CI mirrors for DFlash and DSpark, including NVFP4 targets through ROCm emulation. - Make prompt selection deterministic, remove unnecessary CI progress output, and select platform-appropriate backends. - Strengthen acceptance thresholds, expected-failure handling, and failure diagn…
  • 8a9f9f7 #51735 — [CI] Parallelize release image publishing (#51735)
    • 作者: Kevin H. Luu | +178/-141 | 2 个文件
    • replace the single serial DockerHub publish job with a seven-way Buildkite matrix - publish CUDA 13.0, CUDA 12.9, both Ubuntu 24.04 CUDA variants, ROCm, XPU, and CPU image families independently - keep the existing approval block and preserve all as the script’s backward-compatible default ## Why Release v0.27.0 build 4962 spent more than two hours in the DockerHub publish step. The script seria…
  • 3c6b2e9 #51732 — [CI] Add /ci cancel command (#51732)
    • 作者: Kevin H. Luu | +151/-6 | 4 个文件
    • add an exact /ci cancel PR comment command alongside /ci run and /ci retry - find Buildkite builds in cancelable states for the PR’s current branch, verify each build belongs to that PR, and request cancellation through the Builds API - report canceled build links or a clean no-op when no matching build is active - document the command in the existing CI authorization and first-time contributor …
  • bd65360 #51213 — [XPU][Test] Pin block size in test_multi_connector (#51213)
    • 作者: liuzhenwei | +2/-1 | 2 个文件

    test_multi_example_connector_consistency hardcodes assertions for num_blocks=[7] and 96 external matched tokens in update_state_after_alloc. These expected values rely on a KV cache block_size of 16. While block_size defaults to 16 on CUDA, it defaults to 64 on XPU, causing the test assertions to fail on XPU devices. This PR explicitly passes block_size=16 to fix this issue. pytest -s -v tests/v1/…

  • ea3115e #51557 — [CI] Stabilize DP supervisor lifecycle tests (#51557)
    • 作者: Taneem Ibrahim | +22/-10 | 1 个文件

    Stabilize DP-supervisor lifecycle tests on loaded CI. The readiness helper silently exhausted a 10-second wait, and intentional shutdown rejected connection resets. No open PR fixes these helpers; #50916 only changed signal handling. ## Reproducer Main build 83054: test_basic_lifecycle and test_failed_startup failed. ## Output on main / on branch Main: 2 failed, 253 passed; branch: 16 passed. .ven…

  • 789c4f9 #51566 — [CI] Bump CUTLASS DSL to 4.6.2 (#51566)
    • 作者: Lucas Wilkinson | +2/-2 | 1 个文件
    • bump nvidia-cutlass-dsl[cu13] from 4.6.0 to 4.6.2 - bump quack-kernels from 0.6.1 to 0.6.4, whose package metadata pins CUTLASS DSL 4.6.2 - replace the temporary QuACK downgrade with the CUTLASS DSL fix proposed in #51286 ## Why CUTLASS DSL 4.6.0 rejects the FA4 SM100 split-KV kernel with TYPE_UNSTABLE_JOIN, which can fail B200/B300 attention tests and Kimi-K3 startup during CuTeDSL warmup. CUTL…
  • 51562de #51604 — [CI][XPU] Add VLLM_DISABLE_COMPILE_CACHE=1 for other random failed cases in Intel GPU CI (#51604)
    • 作者: xiangdong | +6/-4 | 3 个文件

    Following https://github.com/vllm-project/vllm/pull/51337, add VLLM_DISABLE_COMPILE_CACHE=1 for other random failed cases in Intel GPU CI ## Test Result —

  • 74c94b9 #51422 — [CI] Upgrade huggingface-hub to 1.27.0 (#51422)
    • 作者: Andreas Karatzas | +9/-9 | 5 个文件
  • 31cd109 #40958 — [ROCm][CI] Extend ROCm AITER MHA (FA) coverage (#40958)
    • 作者: Andreas Karatzas | +906/-258 | 2 个文件

    This PR consolidates the ROCm AITER flash-attention coverage into one backend-named file: test_rocm_aiter_fa.py. It merges the old direct kernel stress file and a new set of tests so the backend and the unique direct-kernel checks live together. That gives the test the same name as the backend it actually exercises, which makes the tree easier to read. There is no intended kernel behavior change h…

🐛 Bug Fix

  • 5af7c8d #51812 — [Bugfix] Align Qwen GDN gates with speculative tokens (#51812)
    • 作者: Jiangyun Zhu | +6/-2 | 1 个文件

    Fix Qwen GDN speculative decoding when a mixed batch places non-speculative tokens before speculative tokens. mixed_qkv is gathered with spec_token_indx, but the fused recurrent update previously received the unsorted a and b gate tensors. The kernel consumes the first T_spec gate rows, so the gates could belong to different tokens than the gathered Q/K/V rows. Gather a and b with the same indices…

  • 0fb9897 #51819 — [Bugfix][MoE] Support GELU tanh in FlashInfer B12x MoE (#51819)
    • 作者: Andrii Skliar | +6/-1 | 1 个文件

    Adds MoEActivation.GELU_TANH support to the FlashInfer B12x MoE backend so Gemma 4 NVFP4 can use –moe-backend flashinfer_b12x. ## Validation Validated with nvidia/Gemma-4-26B-A4B-NVFP4 on DGX Spark using FlashInfer 0.6.16.post3

  • 87668ab #51768 — [Bugfix] Guard DeepSeek V4 MRV1 piecewise CUDA graphs (#51768)
    • 作者: Woosuk Kwon | +73/-0 | 2 个文件

    Default DeepseekV4ForCausalLM to Model Runner V2 and reject only the known-broken configuration: - DeepSeek V4 - Model Runner V1 - PIECEWISE or FULL_AND_PIECEWISE CUDA graphs The validation runs after CUDA-graph mode resolution. MRV1 therefore remains available with eager/NONE, FULL, or FULL_DECODE_ONLY; other model architectures are unaffected. ## Why #51430 exposed a correctness problem in the l…

  • 12bea3e #51145 — [Bugfix][ROCm] Fix DeepSeek V4 DSpark probabilistic startup (#51145)
    • 作者: Tuukka Sarvi | +2/-0 | 1 个文件

    DeepSeek V4 DSpark uses a full-vocabulary draft model, so draft token ids are already target token ids and no draft-to-target remapping is needed. The CUDA implementation exposes this by defining draft_id_to_target_id = None, but the ROCm platform implementation was missing the same marker. The shared DSpark speculator reads model.draft_id_to_target_id when draft_sample_method=“probabilistic” to d…

  • 78e7fdd #51627 — [Bugfix][CPU] Make the Apple Silicon BF16 probe fall back instead of raising (#51627)
    • 作者: Kush Zingade | +8/-5 | 1 个文件

    Problem In CpuPlatform.supported_dtypes, the macOS arm64 branch probes for BF16 with: macOS does not publish the hw.optional.arm.FEAT_BF16 OID when the CPU does not have the feature. sysctl then prints unknown oid to stderr and exits 1. check_output raises CalledProcessError on a nonzero exit, so it never returns a value other than b"1", and the return [torch.float16, torch.float32] line below …

  • 0fec3d6 #51622 — [Bugfix][KV Offload] Centralize shared mmap cleanup in CPU worker (#51622)
    • 作者: AlexHuang | +220/-13 | 2 个文件

    Fix shared mmap ownership and shutdown ordering in the CPU KV offloading worker. CPUOffloadingWorker creates one set of mmap-backed CPU tensors and passes the same tensors to both the GPU-to-CPU and CPU-to-GPU handlers. Previously, the region cleanup responsibility was attached to one direction, even though both directions could still have asynchronous transfers using the same backing memory. This…

  • ce07118 #51766 — [Bugfix][Core] Preserve Mamba running CoW after external hits (#51766)
    • 作者: Dao007forever | +73/-0 | 2 个文件

    Preserve running-request copy-on-write semantics for Mamba align requests whose externally loaded prefix already provides every block needed by their first continuation. The observed Kimi-K3 geometry is: ## Root cause allocate_external_computed_blocks() can populate the request’s Mamba block table, but when the first continuation stays in the same Mamba block, allocate_new_blocks() returns early w…

  • 65b7662 #51556 — [Bugfix][Frontend] Report Cohere stop sequences correctly (#51556)
    • 作者: cherry77-cloud | +63/-5 | 3 个文件
    • report Cohere v2 stop-sequence termination as STOP_SEQUENCE instead of COMPLETE - preserve the OpenAI-compatible stop_reason through the Cohere streaming state machine - distinguish matched string stop sequences from EOS and integer stop-token IDs - add non-streaming and streaming regression coverage ## Root cause vLLM represents EOS, stop-token, and stop-string termination with finish_reason=“s…
  • 07443be #51482 — [BugFix][Core] free_blocks: restore prepend (LIFO) reuse order when prefix caching is off (#51482)
    • 作者: Ming Huang | +1/-1 | 1 个文件

    What #48017 was described as a pure no-op that skips the pointless LRU hash-split in BlockPool.free_blocks when prefix caching is disabled — but the merged condition (block.block_hash is None and self.enable_caching) routes every freed block through append_n instead of prepend_n on that path. Since allocation pops from the head of FreeKVCacheBlockQueue, this silently flipped block reuse…

  • 419b51b #51721 — [Bugfix][ROCm][CI] Stabilize build context and source caches (#51721)
    • 作者: Andreas Karatzas | +345/-61 | 2 个文件
    • materialize the pinned Git revision into an owned, canonical Docker context instead of mutating the mixed-UID Buildkite checkout or copying its Git history; use stable Git archive metadata for wheel versioning - fetch ROCm Triton kernels once from a commit-pinned, SHA-256-verified archive, reuse the extracted package in both ROCm wheel paths, and fail if the Docker/CMake pins drift - probe cache…
  • 99e62b8 #49519 — [Bugfix][Model Loader] Defer post-load attention weight processing (#49519)
    • 作者: aoshen02 | +105/-24 | 4 个文件

    Align layerwise weight loading and reload with the standard post-load attention lifecycle: - identify deferred attention layers through one shared predicate; - cover AttentionLayerBase implementations with a post-load hook, including pluggable multi-head latent attention implementations; - explicitly cover MMEncoderAttention, which has the same lifecycle but does not inherit AttentionLayerBase; - …

  • f37dff2 #51259 — [Bugfix] Import each packed IPC export once on the consumer side (#51259)
    • 作者: acmore | +59/-4 | 2 个文件

    FIX #51258 torch’s cross-process IPC refcount pairs one reduce_tensor export with one consumer rebuild-release cycle: the export sets the shared counter to 1, releasing the rebuilt tensor decrements it, and the producer can only reclaim a dropped buffer via torch.cuda.ipc_collect() when the counter is exactly 0. packed_ipc_producer exports its staging buffer once and ships the same rebuild args wi…

  • b2506d6 #51461 — [MM][CG][BugFix] Fix Ernie-4.5-VL encoder CG postprocess for multi-path outputs (#51461)
    • 作者: Qiuyang Yue | +3/-1 | 1 个文件

    Fix-forward for the crash that prompted #51263 (revert of #45254), instead of reverting. SupportsEncoderCudaGraph.postprocess_encoder_output now receives outputs: dict[str, torch.Tensor] keyed by encoder path (multi-path graph support), but Ernie4_5_VLMoeForConditionalGeneration’s override still treated the first arg as a single tensor: File “vllm/model_executor/models/ernie45_vl.py”, line 1734, i…

  • 8bcc916 #51727 — [Bugfix] Fix DeepSeek V4/3.2 tokenizer vocab size overcount crashing guided decoding (#51727)
    • 作者: Flora Feng | +0/-11 | 2 个文件

    Fixes #50924 #51467. ### Root cause DeepseekV4Tokenizer.len overcounts the vocabulary. vllm/tokenizers/deepseek_v4.py:81-82 returns tokenizer.vocab_size + len(get_added_vocab()). For DeepSeek-V4-Flash-0731 that’s 128000 + 1283 = 129283 — but 3 of those 1283 added tokens have ids below 128000 (they’re already in the base vocab), so the true token count is 129280, exactly config.vocab_size. The …

  • c3cac8c #49815 — [Bugfix][MiMo] Apply vision attention sinks in the window attention path (#49815)
    • 作者: Almog Tavor | +132/-12 | 3 个文件

    Fixes #47864. MiMoVisionAttention allocates self.sinks only when the block is not in fullatt_block_indexes, which is exactly the set of blocks that run _forward_window_attn. That path never read the parameter, so the sink weights were loaded from the checkpoint and dropped. XiaomiMiMo/MiMo-V2.5 ships visual.blocks.N.attn.sinks for exactly those blocks, 24 of its 28, and none for the full attention…

  • 98a4144 #51097 — [Bugfix] Preserve non-logitproc entry points in tests (#51097)
    • 作者: Tristan Rice | +36/-5 | 2 个文件

    The fork-path logits processor test helper replaces importlib.metadata.entry_points process-wide and previously returned an empty result for every group other than vllm.logits_processors. PyTorch nightly now registers built-in distributed backends through torch.distributed.backends, so forked engine workers could not discover NCCL and failed with Unknown c10d backend type NCCL. Delegate non-logitp…

  • 3f142bd #51296 — [Bugfix] Align deepseek v4 parser thinking default with tokenizer (#51296)
    • 作者: Flora Feng | +57/-12 | 3 个文件

    DeepSeek V4’s tokenizer defaults to thinking mode when neither thinking nor enable_thinking is specified. The parser defaulted to content mode, causing reasoning and DSML tool-call markup to leak into content instead of producing structured reasoning_content and tool_calls. The mismatch was introduced by PR #50580. ### This PR This PR aligns the parser with the tokenizer while preserving explicit …

  • 05f0a80 #49227 — [Bugfix][Structured Output] Mask request stop tokens in xgrammar until grammar terminates (#49227)
    • 作者: Flora Feng | +114/-6 | 7 个文件

    Fixes #42403. With structured outputs (e.g. JSON schema), a model can sample a stop token while the grammar FSM is still mid-object, truncating generation into invalid JSON. On Gemma this is <end_of_turn> (id 106). Root cause: xgrammar only knows the tokenizer’s single eos, but a request’s real stop set is generation_config’s eos list plus any user stop_token_ids. On Gemma-4 that gap is …

  • 355a338 #51602 — [BugFix][SpecDecode] Fix dspark parallel_drafting_token_id init bug (#51602)
    • 作者: wangxiyuan | +7/-2 | 1 个文件

    Fix dspark parallel_drafting_token_id init bug in MRV1. The bug has been fixed in MRV2 already in get_parallel_drafting_token_id function. ## Test Result —

  • 0e2d780 #51682 — [Bugfix][Kimi-K3] Give the AMD packed KDA decode kernel the state-index stride (#51682)
    • 作者: Lyu, Xudong | +18/-4 | 2 个文件

    fused_recurrent_kda_packed_decode in the AMD copy of the vendored KDA kernels indexes state_indices as if it were unit-stride, and rejects anything else up front: The NVIDIA copy of the same kernel already takes a stride_state_indices and only requires the tensor to be one-dimensional. The AMD copy was left behind, and this brings it in line — same parameter name, same relaxed check — so t…

  • 436be94 #51635 — [ROCm][Bugfix] Use TCP store when AITER custom all-reduce is enabled (#51635)
    • 作者: vllmellm | +34/-3 | 3 个文件

    #50999 switched single-node executors from TCP to file:// rendezvous to eliminate startup port races. On ROCm with AITER custom all-reduce enabled, that broke every server start at worker init: AITER’s custom all-reduce asserts the default store is a TCPStore (aiter/dist/device_communicators/custom_all_reduce.py); file:// rendezvous produces a FileStore. AITER hasn’t accepted FileStore upstream, s…

  • 3dafaef #51573 — [Bugfix][Core] Emit –no-{key} for false BooleanOptionalAction flags in YAML config (#51573)
    • 作者: Raj Vijay Firke | +24/-0 | 2 个文件

    Fixes #51401 –config YAML files silently drop false boolean values. For BooleanOptionalAction flags (e.g. –enable-flashinfer-autotune) whose default is None and gets resolved later by optimization-level logic, the user’s explicit false was lost — causing unexpected behavior (OOM in the reported case, as flashinfer autotune warmup ran despite being explicitly disabled). ## Root Cause In FlexibleA…

  • 900d09f #50734 — [Bugfix][Model] Fix Qwen3.5 MTP for text-only checkpoints (#50734)
    • 作者: efschu | +55/-7 | 3 个文件

    What Two gaps that keep –speculative-config ‘{“method”:“mtp”,…}’ from working on Qwen3.5 checkpoints that ship only the text config. ## Details 1. The MTP config override does not know the text-only model types. SpeculativeConfig.hf_config_override matches only qwen3_5 / qwen3_5_moe (vllm/config/speculative.py). #50210 registered qwen3_5_text and qwen3_5_moe_text in _CONFIG_REGISTRY (vll…

  • 7303c66 #48171 — [Bugfix] Fix lfm2 tool parser dropping calls with brackets or newline… (#48171)
    • 作者: Zetian Li - ikun | +1186/-24 | 4 个文件

    [Bugfix] Fix lfm2 tool parser dropping or corrupting recoverable tool calls The lfm2 pythonic tool parser silently drops (or corrupts) tool calls for a range of outputs that real agentic models emit routinely. Each commit fixes one failure class, with the model output that triggered it: | Model output | Before this PR | After this PR | |—|—|—| | command=‘grep -F “]” log.txt’ (bracket in st…

🔩 Misc

  • 5b5eae2 #49444 — [Misc] Enable test_silu_mul_fp8_quant_deep_gemm on XPU (#49444)
    • 作者: pmanczak | +19/-17 | 2 个文件

    persistent_masked_m_silu_mul_quant() called current_platform.get_device_capability().to_int() and asserted it wasn’t None. Device capability is a CUDA/ROCm concept, and XpuPlatform returns None, so the wrapper blew up on XPU instead of falling through to the Triton path. 1. batched_deep_gemm_moe.py - gate the C++ kernel on current_platform.is_cuda() and current_platform.has_device_capability(8…

  • adc1200 #51753 — [Misc] Use VLLMValidationError in scoring input validation (#51753)
    • 作者: Frank | +72/-4 | 2 个文件

    Part of #48227. Migrate four client-caused validation errors in vllm/entrypoints/pooling/scoring/utils.py from raw ValueError to VLLMValidationError: - unsupported multimodal input - incompatible input lengths - empty first input - empty second input This preserves the existing validation behavior and error messages. The parameter and value fields remain unset because this utility is shared by /sc…

  • ec21f61 #51672 — [Misc] Enable test_fused_moe_wn16 on XPU (#51672)
    • 作者: pmanczak | +23/-13 | 1 个文件

    Enables XPU coverage for test_fused_moe_wn16, which exercises the fused_moe_kernel_gptq_awq Triton kernel (fused MoE with GPTQ/AWQ INT4/INT8 weight-only quantization). - Hardcoded device=“cuda” replaced with the DEVICE_TYPE already defined in this module (current_platform.device_type). - Added a skipif guard so platforms without the kernel skip instead of failing on a missing device. Test-only dev…

📦 Other

  • 4988df2 #51726 — [Config] Update default _max_num_batched_tokens from 8192 to 16384 (#51726)
    • 作者: Wentao Ye | +12/-31 | 2 个文件

    Update this num to a larger num when gpu memory is enough, perf can be seen https://github.com/vllm-project/vllm/pull/51725. Note that SGLang use the same num I make the PRs separate so easier to review and revert if any issue bumps up.

  • dd7cc85 #51407 — Add MoE output contract for MoE tail fusion (#51407)
    • 作者: Jee Jee Li | +127/-13 | 4 个文件

    Test Result —

  • b64a270 #51447 — Bound generation inputs before expensive work (#51447)
    • 作者: Clinton Thomas | +421/-37 | 11 个文件

    Bound generation inputs before expensive work ## What this fixes Five ordinary request values could make vLLM perform work far larger than the JSON body before an effective limit ran. Examples included tokenizing 50,000 bad words before a later rejection, scanning 500,000 stop strings after every generated token, and scanning an entire DeepSeek message history again for every message. Before thi…

  • 5b184f7 #51308 — connects vLLM Recipes with vLLM’s native config-based deployment and benchmark (#51308)
    • 作者: Louie Tsai | +773/-43 | 6 个文件

    This PR connects vLLM Recipes with vLLM’s native config-based deployment flow. It adds a tool that converts a hardware-specific Recipe into: * config.yml — used by vllm serve –config * env.sh — Recipe-specific environment variables The docs are updated to make this workflow discoverable from CPU model support, deployment, benchmarking, and server configuration pages. ### Flow ### Example See …

  • f863387 #51770 — [XPU] Fix UVA weight offloading (non-pinned-tensor views and static Triton launcher) (#51770)
    • 作者: Chaojun Zhang | +33/-1 | 2 个文件

    Fix two startup crashes when using weight offloading (–cpu-offload-gb) on XPU. ### Issue 1: Pinned-memory assertion failure The assertion fails for zero‑element tensors (common in quantized checkpoints) or when pinning is disabled via VLLM_WEIGHT_OFFLOADING_DISABLE_PIN_MEMORY=1. The underlying XPU kernel already handles both cases, but the Python wrapper enforces the assertion prematurely. Fix: D…

  • d8f8400 #51780 — [Model] Enable tower and connector LoRA for Keye (#51780)
    • 作者: liushujia122 | +8/-0 | 1 个文件
    • Implement get_num_mm_encoder_tokens() and get_num_mm_connector_tokens() on BaseKeyeModule. - Convert between the post-merge multimodal token count seen by the language model and the pre-merge activation length used by the vision tower and Keye projector. Part of #31479. ## Implementation Keye reports language-model image tokens after spatial merging: The vision tower operates on the unmerged pat…
  • 490259c #50977 — profiler: add PrivateUse1 activity support for custom backends (#50977)
    • 作者: David Holtz | +2/-1 | 1 个文件

    This change enables profiler support of vLLM for custom or third-party backends using PyTorch’s PrivateUse1 device / activity namespace. For example, once integrated in spyre-inference this enables profiling on the IBM’s spyre device. n/a ## Test Result n/a —

  • 52c70b2 #50826 — [XPU] [Linear] enable torch linear backend for blockwise gemm on xpu (#50826)
    • 作者: zofia | +2/-0 | 2 个文件

    This pr enables the torch backend on xpu.

  • a311916 #46849 — [MRV2][Spec] Fuse AR speculator multi-step decodes back into one CUDA graph (#46849)
    • 作者: Yizhou | +492/-49 | 10 个文件

    This PR restores fused multi-step CUDA graph execution for autoregressive speculative decoding in Model Runner V2. #41162 fixed stale attention metadata by rebuilding it and replaying a separate CUDA graph for every draft step. While correct, that design reintroduces per-step Python dispatch, metadata construction, and CUDA graph launch overhead. This PR instead captures the post-prefill draft loo…

  • 1044d19 #51473 — [ROCm][DSV4] Preserve native MXFP4 TP8 shard allocation (#51473)
    • 作者: Fangzhou Ai | +63/-3 | 2 个文件

    DeepSeek-V4 has a 3072-wide routed-expert intermediate, so effective MoE TP8 owns a native 384-wide MXFP4 shard. The generic ROCm allocation path rounded that shard to 512 before parameter allocation and checkpoint loading. This was physical padding rather than failed checkpoint sharding, and it inflated the routed-expert allocation by 33%. The rounding rule now lives in the generic MXFP4 oracle b…

  • be2e274 #51773 — Fix docs on main (#51773)
    • 作者: Harry Mellor | +2/-8 | 1 个文件

    The docstring was not updated with the spec of the method.

  • 2acb055 #44201 — [CPU][Zen] Route BF16 MoE inference through zentorch on AMD (#44201)
    • 作者: Priyansh Jain | +316/-3 | 4 个文件

    Routes CPU MoE on AMD Zen through a zentorch-backed fused MoE path in CPUFusedMOE, ahead of the existing AMX grouped-GEMM / OneDNN / per-expert PyTorch fallbacks: - forward_zentorch — MoE FFN via torch.ops.zentorch.zentorch_fused_moe (standard [E, …] expert weights; no cpu_prepack_moe_weight). - is_zentorch_moe_supported() — capability check at CPUFusedMOE init: op registered, moe_config…

  • b76cf75 #50831 — [XPU] install xpu-manager for device monitor (#50831)
    • 作者: Yan Ma | +19/-0 | 1 个文件

    This PR adds xpu-manager installation in vllm XPU image, to help monitor device status. ## Test Result —

  • 608c124 #51733 — [Attention] Fix MLA prefill workspace allocation size (#51733)
    • 作者: Wei Zhao | +7/-11 | 3 个文件

    With #50613, MLA prefill workspace no longer requires to be at least max-num-seqs * block_size. This requirement is added back in https://github.com/vllm-project/vllm/pull/50484. This PR reverts it. - Test test_mla_backends.py and test_sparse_mla_backends.py. - Test Kimi K3 DCP 8 on B300 GSM8k ## Test Result —

  • dc5101f #51734 — replace batch_norm to numerically identical without cudnn (#51734)
    • 作者: Khushali Desai | +80/-19 | 3 个文件

    Fixes #51717 With –mm-device-do-normalize (added in #50411), image normalization runs on device through FusedInputNorm. That module implements the per-channel affine output = (input * rescale_factor - mean) / std by calling F.batch_norm with running_mean=0, running_var=1, eps=0, weight=1/std, bias=-mean/std. On CUDA, F.batch_norm dispatches to cuDNN’s batch-norm kernels, whose batch dimension is …

  • 90fd4a3 #51446 — Preserve revision pins in secondary artifact loaders (#51446)
    • 作者: Clinton Thomas | +226/-13 | 11 个文件

    Preserve revision pins in secondary artifact loaders ## What this fixes An operator can use –revision and –code-revision to select reviewed model artifacts. Several secondary loaders honored the primary model selection but fetched related configuration, processor, quantization, or export files from the repository’s changing default branch. For example, a deployment could pin –revision audited…

  • 0f2ea97 #51688 — [KV Connector][Offloading] Keep per-layer KV registration when canonical_layout is requested (#51688)
    • 作者: Itay Etelis | +31/-1 | 2 个文件

    The connector prefers cross-layer blocks, but a cross-layer slab has no per-layer refs to certify, so uniform-attention models on the V1 model runner refused canonical_layout at startup (e.g. Qwen3-30B-A3B at tp2). Keep per-layer registration when the canonical layout is requested; the preference is unchanged otherwise. Unlocks canonical offload for uniform-attention MoE models on the V1 runner: Q…

  • 48bada6 #51654 — Fix chat completion 500 on non-object JSON bodies (#51654)
    • 作者: Tarun Kumar | +58/-7 | 2 个文件

    While performing contract (negative) testing, it was observed that the server returns an HTTP 500 Internal Server Error instead of an appropriate 4xx client error for invalid request payloads. The vLLM logs also report the failure as an Internal Server Error. Please find the fix detail below - Guard mode=“before” chat-completion validators with isinstance(data, dict) checks so non-object JSON bodi…

  • f1e921b #51144 — [Rust Frontend] Support dynamic tools from developer messages (#51144)
    • 作者: Bugen Zhao | +686/-258 | 28 个文件

    The Rust frontend already carries message-scoped function tools on developer messages, but tool choice resolution, parser activation, and structural-tag generation only used request-level tools. Dynamic-only declarations therefore reached some renderers without becoming available to downstream tool handling. This PR adds a request-scoped ResolvedToolContext that resolves request-level and develope…

  • d8c70f2 #51235 — [Rust Frontend] Upgrade MiniJinja to 2.22 & remove method lookup workaround (#51235)
    • 作者: Bugen Zhao | +12/-98 | 5 个文件

    Upgrade MiniJinja and minijinja-contrib from 2.18.0 to 2.22.0. MiniJinja 2.22 fixes method lookup precedence for mappings with keys such as items, as described in mitsuhiko/minijinja#903. This lets the Rust frontend remove its custom TemplateMap (introduced in #44311) and pass tool definitions and tool-call arguments to MiniJinja as ordinary serde_json::Value values. The upgrade also adopts Jinja2…

  • 3e174bb #47352 — [Model Runner V2][MTP] Share topk index buffer between draft steps (#47352)
    • 作者: Giancarlo Delfin | +68/-6 | 2 个文件

    Context Deepseek-style topk index selection for MTP has a bug during proposal stage when the topk indices are shared among all MTP draft step forward passes. After the first draft step, set_skip_topk is called on the MTP model to force the remaining draft steps to reuse the same topk indices values (held within topk_indices_buffer of the attention module). This feature was introduce in https://g…

  • fa722b9 #51178 — [Rust Frontend][gRPC] Add explicit data-parallel rank routing (#51178)
    • 作者: Connor Carpenter | +215/-60 | 19 个文件

    Add explicit data-parallel rank routing to the Rust frontend gRPC inference API. - Add optional data_parallel_rank to GenerateRequest. - Forward the requested global rank through request conversion into EngineCore selection. - Route only to a connected engine with the matching global identity; retain load balancing when the field is unset. - Reject ranks that cannot be represented by vLLM’s two-by…

  • 635dd6a #51276 — [Build][gRPC] Publish protobuf schemas to Buf (#51276)
    • 作者: Connor Carpenter | +85/-0 | 4 个文件

    Publish vLLM’s canonical gRPC schemas to the Buf Schema Registry (BSR). - Add a named public Buf module rooted at rust/proto. - Build and lint schema changes on pull requests. - Publish the latest Git main schema daily to the Buf nightly label, with manual retry support. - Update the Buf main label and add the matching vX.Y.Z label when vLLM publishes a release tag. - Document the schema source of…

  • 63ac04a #50484 — [Kimi-K3] DCP support (#50484)
    • 作者: Summer Yang | +3529/-82 | 19 个文件

    Co-authored-by: @foraxe Follow-up to #50000 adding decode context parallelism for Kimi-K3: - direct symmetric-memory DCP A2A output/LSE reduction, including empty-KV-shard masking - DCP support for the fused Kimi-K3 MLA layer - NVLS-multicast direct query gather and multimem chunked-context KV gather - direct publication of query shards into the consumer-final buffer, removing the staging-to-final…

  • d40c3e3 #50693 — Fix DSpark warmup without sparse index buffer (#50693)
    • 作者: xijiaat | +5/-2 | 1 个文件

    Fixes #50615. During DSpark startup, profile_run calls forward_mqa without attention metadata to size the warmup workspace. For the SWA-only draft layer (compress_ratio <= 1), topk_indices_buffer is intentionally not allocated, but the warmup path asserted that the buffer existed before setting top_k to zero. This makes the engine fail after loading the full target and draft model. This change han…

  • 1482d2e #47030 — [ROCm][DistInf] Enable vLLM DI CI with buildkite/slurm (#47030)
    • 作者: Chaitanya Sri Krishna Lolla | +1974/-0 | 6 个文件

    This PR enables the Buildkite CI scirpts that exercises vLLM Disaggregated inference P/D with MoRI IO KV connector on ROCm AMD devices. This brings up a full PD topology health-gates on everyserver, runs the GSM8k accuracy. This PR supports with two modes WIDE_EP_MODE=0 / 1 to support general PD disaggregation wtih TP8 indepdent servers (When WIDE_EP_MODE=0) and DP + Expert Parallelism (when WIDE_…

  • 21c667a #51184 — [Docker] Cache test dependencies before vLLM install (#51184)
    • 作者: Michael Goin | +61/-6 | 4 个文件
    • split the runtime-only setup into a shared vllm-runtime-base stage - add a test-deps stage that installs Git and requirements/dev.txt before per-commit artifacts - install the vLLM and EP wheels, then add the repository source, in the final test stage - regenerate and document the Docker stage graph, concentrating parallel edges for readability The production vllm-base path still uses the same v…
  • 3a749ce #51424 — [Build] Skip precompiled wheel fetch during metadata hooks (#51424)
    • 作者: Michael Goin | +6/-2 | 1 个文件

    VLLM_USE_PRECOMPILED=1 uv pip install –editable . invokes setup.py for egg_info, dist_info, and the actual editable build. Because precompiled-wheel setup ran at module scope, each phase could resolve, download, and extract the same wheel. Skip precompiled artifact setup for the two metadata-only commands. The actual wheel build path is unchanged. This does not duplicate an open PR. I searched op…

  • 640a090 #49948 — Fix DoS via sample-rate forgery bypassing audio decode duration guard (#49948)
    • 作者: Juan Pérez de Algaba | +126/-7 | 5 个文件

    The existing max_duration_s guard in load_audio_soundfile computes duration as f.frames / f.samplerate, trusting the container header. An attacker can set samplerate=655350 (FLAC max) with 8 channels to make hundreds of millions of frames appear as a short clip, bypassing the duration check while f.read() allocates up to 11.7 GiB of float32 PCM — enough to OOM-kill the API server. Add VLLM_MAX_AUD…

  • 37fbf52 #51011 — [ROCm][MLA] [K3] Fix fp8 KV cache decode on the AITER MLA backend (#51011)
    • 作者: fanxingran | +366/-63 | 4 个文件

    Kimi-K3 at TP8 has 12 MLA heads per rank and cannot serve correctly with –kv-cache-dtype fp8 on the ROCm AITER MLA backend. On unmodified main, a full GSM8K run at that configuration scores 74.00% with 285 of 1319 answers degenerate; with this PR it scores 97.19% with none. The backend fails three different ways depending on head count and query length, and the third is the danger…

  • 8fea1d3 #43680 — Fix uniform_random routing simulation to sample without replacement (#43680)
    • 作者: Elvir Crnčević | +4/-8 | 1 个文件

    torch.randint can pick the same expert twice for a single token, violating DeepEP v2’s dispatch kernel precondition that topk selections within a warp are deduplicated. Use torch.topk on random scores instead, which inherently produces distinct indices.

  • cf8f3a3 #51657 — [2/N] Harden Transformers modelling backend multi-modal path (#51657)
    • 作者: Harry Mellor | +276/-136 | 3 个文件

    Continues from https://github.com/vllm-project/vllm/pull/51408 ## Added - –mm-encoder-only and –limit-mm-per-prompt =0 now skip weights for this backend, via a _mark_model_components hook around AutoModel.from_config - Audio encoders can be torch compiled; compile_mm_encoder was hardcoded to the image encoder ## Fixed - _get_prompt_updates returned None against a declared Sequence[Prom…

  • 243c63b #51603 — [V1][Scheduler] Apply Mamba alignment before encoder caps (#51603)
    • 作者: Jiangyun Zhu | +105/-20 | 2 个文件
    • Apply Mamba block alignment before multimodal encoder scheduling can cap a prefill chunk. - Use EAGLE’s shifted encoder window consistently when calculating the overlapping embedding range. - Preserve the existing align-mode policy that waits for a fresh scheduler budget when a complete Mamba block normally fits in one step. - Add a scheduler-level regression for two images that each fit the enc…
  • d89ba64 #50441 — [XPU] bump up xpu kernel to v0.1.12.3 (#50441)
    • 作者: Kunshang Ji | +1/-1 | 1 个文件

    use pypi for vllm-xpu-kernels dependency https://pypi.org/project/vllm-xpu-kernels ## Test Result —

  • 70b84f0 #49797 — Fix Gemma 4 for upcoming Transformers version (#49797)
    • 作者: Harry Mellor | +570/-106 | 15 个文件

    Transformers v5.15.0 will introduce heterogeneous config machinery that has been adopted by Gemma 4. This PR updates the Gemma 4 implementations and the Transformers modelling backend to be compatible with these new heterogeneous configs. It also required making some small changes to the model arch converter so that it plays nicely with heterogeneous configs. Supersedes https://github.com/vllm-pro…

  • 81840a1 #48414 — [KV Connector] Canonical CPU layout for parallelism-agnostic KV offload (#48414)
    • 作者: Itay Etelis | +891/-103 | 12 个文件

    Stacked on #48408. Stores offloaded KV in the canonical (parallelism-free) layout described by the refs’ page mappings: each worker scatters its page fragments to their canonical positions in a CPU area shared by the whole worker group. MLA latent and replicated GQA heads are stored once instead of once per rank (empty store runs on non-writers). Copy expansions are precomputed per ref at init; pe…

✨ New Feature

  • 6c95a64 #49315 — [2/N][Feat][Perf] Add new warmup infrastructure for JITs. Add predicate filtering for JIT warmup, and migrate Inkling FA4 (#49315)
    • 作者: Roberto L. Castro | +743/-390 | 12 个文件

    Description This PR migrates Inkling FA4 attention warmup to the shared JIT warmup contract introduced in #47451, and deprecates the legacy CuTeDSL warmup path (there is a conflict with the Kimi K3 integration #50089 that prevents this deprecation from being completed. It will be addressed in a future PR). See https://github.com/vllm-project/vllm/issues/49349 for more context ### Motivation …

📖 Documentation

  • 529d010 #51500 — [Doc] Fix typos in speculative decoding docs (#51500)
    • 作者: Kyungmin Lee | +8/-8 | 4 个文件

    This PR fixes typos in the speculative decoding documentation. ## Test Result —

  • 3a79957 #49353 — [Doc] Add Crusoe Managed Inference deployment guide (#49353)
    • 作者: Emmanuel Acheampong | +64/-0 | 1 个文件

    Adds a deployment guide for Crusoe Managed Inference, an OpenAI-compatible API powered by vLLM. This is a refresh of #36935, which the stale bot closed before it got a review and GitHub wouldn’t let me reopen. Compared to that version this one: - Updates the API endpoint to the current api.inference.crusoecloud.com - Restructures the page to match the other framework docs like runpod.md and dstack…

🖥️ Kernel

  • c76a425 #51739 — [Kernel] Optimize long-context MLA cache gathers (#51739)
    • 作者: Yongye Zhu | +810/-235 | 6 个文件
    • schedule MLA cache gathers by logical cache page instead of independently mapping every output token - flatten page copies/conversions to generate coalesced vector loads and stores, including partial first/last pages and nonzero seq_starts - coalesce FP8-to-BF16 stores and tune CTA geometry for long, uneven prefill workloads - add long-context correctness coverage and a reusable kernel benchmark…

🦀 Rust Frontend

  • 3358490 #51478 — [Frontend] Add content_parts to /inference/v1/generate for raw multim… (#51478)
    • 作者: aoshen02 | +150/-7 | 7 个文件

    …odal input Add a content_parts field to the generate request that accepts OpenAI-style content parts (image_url, input_audio, video_url, etc.). The generate server resolves media internally — no pixel data transfer needed. This enables RL frameworks to pass token_ids + raw media in a single request without going through chat/completions or the render pipeline. Both Rust and Python frontends are u…