共 38 个 commit,涉及 167 个文件,+9449/-2226 行变动。

概要

统计项数值
Commit 数38
变更文件167
新增行数+9449
删除行数-2226

Commit 列表

📦 Other

  • 640a090 #49948 — Fix DoS via sample-rate forgery bypassing audio decode duration guard (#49948)
    • 作者: Juan Pérez de Algaba | +126/-7 | 5 个文件

    The existing max_duration_s guard in load_audio_soundfile computes duration as f.frames / f.samplerate, trusting the container header. An attacker can set samplerate=655350 (FLAC max) with 8 channels to make hundreds of millions of frames appear as a short clip, bypassing the duration check while f.read() allocates up to 11.7 GiB of float32 PCM — enough to OOM-kill the API server. Add VLLM_MAX_AUD…

  • 37fbf52 #51011 — [ROCm][MLA] [K3] Fix fp8 KV cache decode on the AITER MLA backend (#51011)
    • 作者: fanxingran | +366/-63 | 4 个文件

    Kimi-K3 at TP8 has 12 MLA heads per rank and cannot serve correctly with –kv-cache-dtype fp8 on the ROCm AITER MLA backend. On unmodified main, a full GSM8K run at that configuration scores 74.00% with 285 of 1319 answers degenerate; with this PR it scores 97.19% with none. The backend fails three different ways depending on head count and query length, and the third is the danger…

  • 8fea1d3 #43680 — Fix uniform_random routing simulation to sample without replacement (#43680)
    • 作者: Elvir Crnčević | +4/-8 | 1 个文件

    torch.randint can pick the same expert twice for a single token, violating DeepEP v2’s dispatch kernel precondition that topk selections within a warp are deduplicated. Use torch.topk on random scores instead, which inherently produces distinct indices.

  • cf8f3a3 #51657 — [2/N] Harden Transformers modelling backend multi-modal path (#51657)
    • 作者: Harry Mellor | +276/-136 | 3 个文件

    Continues from https://github.com/vllm-project/vllm/pull/51408 ## Added - –mm-encoder-only and –limit-mm-per-prompt =0 now skip weights for this backend, via a _mark_model_components hook around AutoModel.from_config - Audio encoders can be torch compiled; compile_mm_encoder was hardcoded to the image encoder ## Fixed - _get_prompt_updates returned None against a declared Sequence[Prom…

  • 243c63b #51603 — [V1][Scheduler] Apply Mamba alignment before encoder caps (#51603)
    • 作者: Jiangyun Zhu | +105/-20 | 2 个文件
    • Apply Mamba block alignment before multimodal encoder scheduling can cap a prefill chunk. - Use EAGLE’s shifted encoder window consistently when calculating the overlapping embedding range. - Preserve the existing align-mode policy that waits for a fresh scheduler budget when a complete Mamba block normally fits in one step. - Add a scheduler-level regression for two images that each fit the enc…
  • d89ba64 #50441 — [XPU] bump up xpu kernel to v0.1.12.3 (#50441)
    • 作者: Kunshang Ji | +1/-1 | 1 个文件

    use pypi for vllm-xpu-kernels dependency https://pypi.org/project/vllm-xpu-kernels ## Test Result —

  • 70b84f0 #49797 — Fix Gemma 4 for upcoming Transformers version (#49797)
    • 作者: Harry Mellor | +570/-106 | 15 个文件

    Transformers v5.15.0 will introduce heterogeneous config machinery that has been adopted by Gemma 4. This PR updates the Gemma 4 implementations and the Transformers modelling backend to be compatible with these new heterogeneous configs. It also required making some small changes to the model arch converter so that it plays nicely with heterogeneous configs. Supersedes https://github.com/vllm-pro…

  • 81840a1 #48414 — [KV Connector] Canonical CPU layout for parallelism-agnostic KV offload (#48414)
    • 作者: Itay Etelis | +891/-103 | 12 个文件

    Stacked on #48408. Stores offloaded KV in the canonical (parallelism-free) layout described by the refs’ page mappings: each worker scatters its page fragments to their canonical positions in a CPU area shared by the whole worker group. MLA latent and replicated GQA heads are stored once instead of once per rank (empty store runs on non-writers). Copy expansions are precomputed per ref at init; pe…

  • 61c1dd0 #49579 — [EC Connector] Call to EC Connector update_connector_output from scheduler (#49579)
    • 作者: omerpaz95 | +43/-0 | 3 个文件

    The scheduler builds EC (Encoder Cache) connector metadata and calls update_state_after_alloc, request_finished, etc. on the EC Connector, but update_connector_output — the hook that lets an EC connector consume worker-side ECConnectorOutput — was never invoked from anywhere. This wires it into Scheduler.update_from_output, mirroring the existing KV connector call (self.connector.update_connector_…

  • ba1cdcf #51265 — [Model][Quantization] Add Ling-3.0-flash-fp8 support (#51265)
    • 作者: zexplorerhj | +298/-36 | 6 个文件

    This PR follows up on #51045, which added Ling 3.0 Flash BF16, MTP, reasoning-parser, and tool-parser support. It adds loading and serving support for the serialized block-FP8 Ling 3.0 Flash checkpoint while preserving the existing BF16 and non-block-FP8 behavior. ## Why the exclusion matching change is shared Ling/Bailing differs from the checkpoints handled by the current exact matcher because i…

  • b22afe4 #51148 — [CPU] Enable GPTQ and AWQ quantization for s390x (#51148)
    • 作者: Rehan Khan | +88/-24 | 5 个文件

    Enable GPTQ and AWQ quantization for s390x Run inference server with both quantization models and check the result ## Test Result AWQ Model GPTQ Model —

  • 7ce84b9 #51379 — [CPU] Restore linear dispatch for small unquantized GEMMs (#51379)
    • 作者: Li, Jiang | +65/-6 | 1 个文件

    PR #50801 removed the SGLang AMX weight_packed_linear path from the CPU unquantized (bf16/fp16) linear dispatch, leaving oneDNN’s onednn_mm as the only kernel for that path. For small weights – most notably MoE router/gate projections, where N is the expert count rather than a hidden-size-scaled dimension – oneDNN never reaches its compute-bound regime regardless of batch size, so the SGL kernel…

  • a123159 #48798 — Add tiering offloading metrics (#48798)
    • 作者: Srinivas Krovvidi | +1031/-237 | 16 个文件
    • Covers the scope of TieringOffloadingSpec-Level Metrics in #44008 - add TieringOffloadingSpec-level metric definitions with a per-tier label - extend secondary-tier JobResult payloads with optional transfer size/time - emit tiering read/write, failure, and block query/hit counters from TieringOffloadingManager - cover metric registration and manager aggregation paths in focused tests
  • 04d13b5 #51529 — [K3] Allow tpu to import kimi_k3.common (#51529)
    • 作者: Jeff (Junze) Ma | +9/-2 | 1 个文件

    TPU platform plugins reuse Kimi K3’s shared multimodal preprocessing while registering their own model implementations. Avoid eagerly importing a GPU implementation when the active platform’s device type is TPU. CUDA and ROCm model selection remains unchanged. ## Not a duplicate PR #51196 disables dynamic torch.compile for the Kimi vision encoder on TPU; it does not address eager GPU model imports…

  • c423998 #51243 — [KV Offload] Emit self-describing events for partial recurrent blocks (#51243)
    • 作者: Chauncey | +181/-6 | 5 个文件

    Emit self-describing CPU-offload KV events for full-attention cache groups, including hash-aligned metadata for a partial recurrent tail. Mamba/SSM groups intentionally keep the legacy placeholder CPU-offload payload. Their placeholder BlockStored events contain the offload-key hash but have token_ids=[], block_size=0, and no cache-spec metadata. Result: both hooks passed. ### E2E Server: The even…

  • f18e10a #50892 — Bump Flashinfer version to 0.6.16.post3 (#50892)
    • 作者: Wei Zhao | +51/-58 | 5 个文件

    Bump flashinfer version to 0.6.16.post1 and re-enable persistent cache for autotuning using set_autotune_process_group, which supports distributed auto-tuning across ranks with synchronization to avoid timeout caused by straggler. ## Test Result —

  • 1b0ce31 #49328 — [KV Offload] Fix failed-load livelock by marking the lookup verdict as a miss (#49328)
    • 作者: Robbie J | +488/-45 | 9 个文件

    Fixes #49176. A failed secondary-tier load can wedge a request forever. When a promotion (secondary → primary load) fails because the block file is truncated, deleted, or unreadable, the async lookup cache keeps the positive verdict it recorded earlier and never learns about the failure. The scheduler sees a hit, schedules the promotion, the load fails, and next step it sees the same hit again: a …

⚡ Performance

  • fac808b #49436 — [Perf][Hybrid] 3D-grid tiling of the state-copy Triton kernels (#49436)
    • 作者: Francesco Fusco | +347/-68 | 3 个文件

    Follow-up to https://github.com/vllm-project/vllm/pull/48110. This PR further improves the performance by adding a 3D grid and lift the hard 8B alignment precondition by using a head/body/tail pattern. Those are the main changes: 1. Tiles the temporal copy across CTAs. The uint64 body range of each temporal state is partitioned across TEMPORAL_TILES CTAs along a new program_id(2) axis. Grid go…

🧪 CI/Tests

  • bd65360 #51213 — [XPU][Test] Pin block size in test_multi_connector (#51213)
    • 作者: liuzhenwei | +2/-1 | 2 个文件

    test_multi_example_connector_consistency hardcodes assertions for num_blocks=[7] and 96 external matched tokens in update_state_after_alloc. These expected values rely on a KV cache block_size of 16. While block_size defaults to 16 on CUDA, it defaults to 64 on XPU, causing the test assertions to fail on XPU devices. This PR explicitly passes block_size=16 to fix this issue. pytest -s -v tests/v1/…

  • ea3115e #51557 — [CI] Stabilize DP supervisor lifecycle tests (#51557)
    • 作者: Taneem Ibrahim | +22/-10 | 1 个文件

    Stabilize DP-supervisor lifecycle tests on loaded CI. The readiness helper silently exhausted a 10-second wait, and intentional shutdown rejected connection resets. No open PR fixes these helpers; #50916 only changed signal handling. ## Reproducer Main build 83054: test_basic_lifecycle and test_failed_startup failed. ## Output on main / on branch Main: 2 failed, 253 passed; branch: 16 passed. .ven…

  • 789c4f9 #51566 — [CI] Bump CUTLASS DSL to 4.6.2 (#51566)
    • 作者: Lucas Wilkinson | +2/-2 | 1 个文件
    • bump nvidia-cutlass-dsl[cu13] from 4.6.0 to 4.6.2 - bump quack-kernels from 0.6.1 to 0.6.4, whose package metadata pins CUTLASS DSL 4.6.2 - replace the temporary QuACK downgrade with the CUTLASS DSL fix proposed in #51286 ## Why CUTLASS DSL 4.6.0 rejects the FA4 SM100 split-KV kernel with TYPE_UNSTABLE_JOIN, which can fail B200/B300 attention tests and Kimi-K3 startup during CuTeDSL warmup. CUTL…
  • 51562de #51604 — [CI][XPU] Add VLLM_DISABLE_COMPILE_CACHE=1 for other random failed cases in Intel GPU CI (#51604)
    • 作者: xiangdong | +6/-4 | 3 个文件

    Following https://github.com/vllm-project/vllm/pull/51337, add VLLM_DISABLE_COMPILE_CACHE=1 for other random failed cases in Intel GPU CI ## Test Result —

  • 74c94b9 #51422 — [CI] Upgrade huggingface-hub to 1.27.0 (#51422)
    • 作者: Andreas Karatzas | +9/-9 | 5 个文件
  • 31cd109 #40958 — [ROCm][CI] Extend ROCm AITER MHA (FA) coverage (#40958)
    • 作者: Andreas Karatzas | +906/-258 | 2 个文件

    This PR consolidates the ROCm AITER flash-attention coverage into one backend-named file: test_rocm_aiter_fa.py. It merges the old direct kernel stress file and a new set of tests so the backend and the unique direct-kernel checks live together. That gives the test the same name as the backend it actually exercises, which makes the tree easier to read. There is no intended kernel behavior change h…

  • 83ad767 #51539 — [CI] fix docs on main (#51539)
    • 作者: Harry Mellor | +2/-2 | 2 个文件

    Fixes:

  • 7f6432c #48646 — [ROCm][CI] Reuse equivalent ROCm CI images (#48646)
    • 作者: Andreas Karatzas | +1530/-907 | 13 个文件
    • split the ROCm base build into independently cacheable dependency stages - reuse base and ci_base images by deterministic content identity, with trust-scoped writes and digest-pinned downstream handoffs - keep csrc and Rust compiler caches stable across PR commits when their real inputs are unchanged - use the Buildkite checkout for both AMD image steps, while keeping remote-source cache identit…

🔩 Misc

  • ec21f61 #51672 — [Misc] Enable test_fused_moe_wn16 on XPU (#51672)
    • 作者: pmanczak | +23/-13 | 1 个文件

    Enables XPU coverage for test_fused_moe_wn16, which exercises the fused_moe_kernel_gptq_awq Triton kernel (fused MoE with GPTQ/AWQ INT4/INT8 weight-only quantization). - Hardcoded device=“cuda” replaced with the DEVICE_TYPE already defined in this module (current_platform.device_type). - Added a skipif guard so platforms without the kernel skip instead of failing on a missing device. Test-only dev…

🐛 Bug Fix

  • 436be94 #51635 — [ROCm][Bugfix] Use TCP store when AITER custom all-reduce is enabled (#51635)
    • 作者: vllmellm | +34/-3 | 3 个文件

    #50999 switched single-node executors from TCP to file:// rendezvous to eliminate startup port races. On ROCm with AITER custom all-reduce enabled, that broke every server start at worker init: AITER’s custom all-reduce asserts the default store is a TCPStore (aiter/dist/device_communicators/custom_all_reduce.py); file:// rendezvous produces a FileStore. AITER hasn’t accepted FileStore upstream, s…

  • 3dafaef #51573 — [Bugfix][Core] Emit –no-{key} for false BooleanOptionalAction flags in YAML config (#51573)
    • 作者: Raj Vijay Firke | +24/-0 | 2 个文件

    Fixes #51401 –config YAML files silently drop false boolean values. For BooleanOptionalAction flags (e.g. –enable-flashinfer-autotune) whose default is None and gets resolved later by optimization-level logic, the user’s explicit false was lost — causing unexpected behavior (OOM in the reported case, as flashinfer autotune warmup ran despite being explicitly disabled). ## Root Cause In FlexibleA…

  • 900d09f #50734 — [Bugfix][Model] Fix Qwen3.5 MTP for text-only checkpoints (#50734)
    • 作者: efschu | +55/-7 | 3 个文件

    What Two gaps that keep –speculative-config ‘{“method”:“mtp”,…}’ from working on Qwen3.5 checkpoints that ship only the text config. ## Details 1. The MTP config override does not know the text-only model types. SpeculativeConfig.hf_config_override matches only qwen3_5 / qwen3_5_moe (vllm/config/speculative.py). #50210 registered qwen3_5_text and qwen3_5_moe_text in _CONFIG_REGISTRY (vll…

  • 7303c66 #48171 — [Bugfix] Fix lfm2 tool parser dropping calls with brackets or newline… (#48171)
    • 作者: Zetian Li - ikun | +1186/-24 | 4 个文件

    [Bugfix] Fix lfm2 tool parser dropping or corrupting recoverable tool calls The lfm2 pythonic tool parser silently drops (or corrupts) tool calls for a range of outputs that real agentic models emit routinely. Each commit fixes one failure class, with the model output that triggered it: | Model output | Before this PR | After this PR | |—|—|—| | command=‘grep -F “]” log.txt’ (bracket in st…

  • 3b4c86e #51419 — [Bugfix][Quantization] Fix fp32 weight scale for mxfp4 quantization and per-expert checkpoint mapping (#51419)
    • 作者: Isotr0py | +67/-0 | 2 个文件
    • Some mxfp4 checkpoints will store weight_scale as FP32, which will cause casting during weight loading into uint8 - Also fix missing weights mapping for per-expert quantized checkpoints. ## Test Result —
  • 11ba93f #50999 — [BugFix] Use file:// rendezvous for single-node executors to eliminate startup port races (#50999)
    • 作者: aoshen02 | +142/-13 | 6 个文件

    Problem MultiprocExecutor and UniProcExecutor used get_open_port() to make their initial torch.distributed rendezvous URI. That helper binds a probe socket, closes it, and returns the port; the real TCPStore bind occurs only after worker startup. A competing local process can bind the released port in that interval: For the local executor rendezvous this TCP listener is unnecessary. ## Fix Gene…

  • 0820125 #50528 — [Bugfix][Parser] Emit REASONING_END for Inkling tool calls that follow no thinking block (#50528)
    • 作者: Jason | +160/-12 | 4 个文件

    Fixes #50512. With –tool-call-parser inkling –reasoning-parser inkling, a streaming turn that opens straight into its tool block emitted zero tool calls and returned the whole <|content_invoke_tool_json|>{…}<|end_message|> markup as delta.content. The issue title says multi-turn; that is not the discriminant. What decides it is **whether the generated turn confirms the reasoning boundary befor…

  • d694130 #50344 — [BugFix] Scope divergent hybrid cache hits to capable connectors (#50344)
    • 作者: Yifan Qiao | +130/-27 | 8 个文件
    • Gate divergent per-group local cache hits behind a connector capability. - Opt NIXL in and keep unknown/store-style connectors on the common local hit. - Enable the capability for MultiConnector only when every child supports it. - Reconcile divergent partial hits before external lookup. ## Why Per-group hybrid cache hits rely on the connector restoring missing Mamba state. Applying that policy …
  • eb24bc3 #51161 — [Bugfix][KV Offload] Handle chunked local attention in offloading scheduler (#51161)
    • 作者: Almog Tavor | +24/-0 | 2 个文件

    Fixes #51150. get_sliding_window_size_in_chunks() handles SlidingWindowSpec and MambaSpec, then asserts everything else is FullAttentionSpec. Llama 4 uses ChunkedLocalAttentionSpec, so enabling the offloading connector on any Llama 4 checkpoint kills the engine at startup with a bare AssertionError. Chunked local attention never attends further back than one attention chunk, so its reachable tail …

📖 Documentation

  • 3a79957 #49353 — [Doc] Add Crusoe Managed Inference deployment guide (#49353)
    • 作者: Emmanuel Acheampong | +64/-0 | 1 个文件

    Adds a deployment guide for Crusoe Managed Inference, an OpenAI-compatible API powered by vLLM. This is a refresh of #36935, which the stale bot closed before it got a review and GitHub wouldn’t let me reopen. Compared to that version this one: - Updates the API endpoint to the current api.inference.crusoecloud.com - Restructures the page to match the other framework docs like runpod.md and dstack…

🖥️ Kernel

  • 751f2cc #47205 — [Kernel][XPU] Tensor-descriptor operand loads for Triton W8A8 scaled_mm (#47205)
    • 作者: Lena Onyshchenko | +121/-8 | 5 个文件

    Two changes to the compressed-tensors W8A8 INT8 linear path. 1. Tensor-descriptor operand loads for scaled_mm_kernel 2. Enables that path on XPU at all. ### Per-kernel microbenchmark: B70, int8 in / bf16 out, device self-time via torch.profiler: TD reduces self-time by 89-99% across M in {1, 64, 256} and K=N in {4096, 8192}, with max|plain - TD| = 0.0. ### Accuracy gsm8k, 5-shot, n=200: | | fl…