共 21 个 commit,涉及 108 个文件,+4136/-3222 行变动。
概要
| 统计项 | 数值 |
|---|---|
| Commit 数 | 21 |
| 变更文件 | 108 |
| 新增行数 | +4136 |
| 删除行数 | -3222 |
Commit 列表
🧪 CI/Tests
- 83ad767 #51539 — [CI] fix docs on
main(#51539)- 作者: Harry Mellor | +2/-2 | 2 个文件
Fixes:
- 7f6432c #48646 — [ROCm][CI] Reuse equivalent ROCm CI images (#48646)
- 作者: Andreas Karatzas | +1530/-907 | 13 个文件
- split the ROCm base build into independently cacheable dependency stages - reuse base and ci_base images by deterministic content identity, with trust-scoped writes and digest-pinned downstream handoffs - keep csrc and Rust compiler caches stable across PR commits when their real inputs are unchanged - use the Buildkite checkout for both AMD image steps, while keeping remote-source cache identit…
- 7581c56 #51457 — [Test] Add ROCm AITER FP8 MLA prefill accuracy test (#51457)
- 作者: Aarushi Jain | +161/-0 | 1 个文件
Adds a gfx950-only (MI355) accuracy test for the AITER FP8 MLA prefill path (AiterMLAImpl._mla_fp8_prefill_attn -> mla_prefill_ps_asm_fwd + mla_reduce_v1), which previously had no coverage. It drives the real metadata builder (get_ps_metadata_v1) and impl via object.new so it exercises the actual persistent-scheduling metadata contract plus both kernels, and compares the output against a causa…
📦 Other
- 04d13b5 #51529 — [K3] Allow tpu to import kimi_k3.common (#51529)
- 作者: Jeff (Junze) Ma | +9/-2 | 1 个文件
TPU platform plugins reuse Kimi K3’s shared multimodal preprocessing while registering their own model implementations. Avoid eagerly importing a GPU implementation when the active platform’s device type is TPU. CUDA and ROCm model selection remains unchanged. ## Not a duplicate PR #51196 disables dynamic torch.compile for the Kimi vision encoder on TPU; it does not address eager GPU model imports…
- c423998 #51243 — [KV Offload] Emit self-describing events for partial recurrent blocks (#51243)
- 作者: Chauncey | +181/-6 | 5 个文件
Emit self-describing CPU-offload KV events for full-attention cache groups, including hash-aligned metadata for a partial recurrent tail. Mamba/SSM groups intentionally keep the legacy placeholder CPU-offload payload. Their placeholder BlockStored events contain the offload-key hash but have token_ids=[], block_size=0, and no cache-spec metadata. Result: both hooks passed. ### E2E Server: The even…
- f18e10a #50892 — Bump Flashinfer version to 0.6.16.post3 (#50892)
- 作者: Wei Zhao | +51/-58 | 5 个文件
Bump flashinfer version to 0.6.16.post1 and re-enable persistent cache for autotuning using set_autotune_process_group, which supports distributed auto-tuning across ranks with synchronization to avoid timeout caused by straggler. ## Test Result —
- 1b0ce31 #49328 — [KV Offload] Fix failed-load livelock by marking the lookup verdict as a miss (#49328)
- 作者: Robbie J | +488/-45 | 9 个文件
Fixes #49176. A failed secondary-tier load can wedge a request forever. When a promotion (secondary → primary load) fails because the block file is truncated, deleted, or unreadable, the async lookup cache keeps the positive verdict it recorded earlier and never learns about the failure. The scheduler sees a hit, schedules the promotion, the load fails, and next step it sees the same hit again: a …
- 073c510 #43529 — [Migration] Migrate bitsandbytes support to OOT plugin (#43529)
- 作者: Isotr0py | +187/-1895 | 31 个文件
After this PR, BNB support will be migrated to https://github.com/vllm-project/vllm-bnb-plugin, you can still use BNB models normally after plugin installation! - Related issues: https://github.com/vllm-project/vllm/issues/39583 ## Test Result All tests passed —
- dedbf6b #48668 — [V1][Metrics] Preserve prefix-cache stats on zero-output steps (#48668)
- 作者: Rishi Puri | +7/-7 | 1 个文件
Scheduler.make_stats() drains the prefix-cache hit/query counters every step and attaches them to a (possibly empty) EngineCoreOutputs. But the offline LLMEngine.step() path only records stats when the step produced request outputs: So prefix-cache counters consumed on a step that emits no request output — e.g. a cache hit applied when a request is admitted on a non-final prefill chunk — are drain…
- cb1a52a #51058 — [Build] Upgrade runtime image to Ubuntu 24.04, pick up rdma-core > 44 (#51058)
- 作者: Tyler Michael Smith | +7/-5 | 3 个文件
I don’t know any blockers to doing this, and it will also simplify our build/release process since we will no longer have to produce separate images for those who need to use 24.04 The impetus now: NCCL GIN (NCCL_GIN_TYPE=3 aka GDA-KI) requires rdma-core > 44. Ubuntu 22.04 has 39. Ubuntu 24.04 gives us 50.0
- fbff187 #51455 — [Core] Make the GPU sync check thread-local and fix its suppressors (#51455)
- 作者: Nick Hill | +288/-120 | 2 个文件
eImprovements to the VLLM_GPU_SYNC_CHECK machinery added in #44800, found by enabling it across CI. torch.cuda.set_sync_debug_mode is process-global, not thread-local. Two consequences, both verified experimentally: a background thread inherits the mode the main thread armed and raises on syncs that are deliberate (this killed the EPLB async transfer worker), and a set_sync_debug_mode(0) on any th…
- 653ebb5 #40116 — Add torch compile for qwen3_vl encoder (#40116)
- 作者: Tianyu Guo | +22/-1 | 1 个文件
torch.compile can accelerate the computation of the encoder ## Test Result —
🐛 Bug Fix
- eb24bc3 #51161 — [Bugfix][KV Offload] Handle chunked local attention in offloading scheduler (#51161)
- 作者: Almog Tavor | +24/-0 | 2 个文件
Fixes #51150. get_sliding_window_size_in_chunks() handles SlidingWindowSpec and MambaSpec, then asserts everything else is FullAttentionSpec. Llama 4 uses ChunkedLocalAttentionSpec, so enabling the offloading connector on any Llama 4 checkpoint kills the engine at startup with a bare AssertionError. Chunked local attention never attends further back than one attention chunk, so its reachable tail …
- f8d03e7 #51185 — [Bugfix][Build] Patch stable string memleak fix from 2.14 for 2.13 (#51185)
- 作者: Jane (Yuan) Xu | +43/-0 | 2 个文件
The stable ABI header had a memory leak for strings that was patched recently (https://github.com/pytorch/pytorch/commit/7f0ec65bfab968887fc9f23bf319ba3413e1f36f) and will make it into our 2.14 release. This PR ports it over for the interim. This PR should be reverted once torch 2.14 is the build version for vllm. Note that this does not change versioning support for the stable ABI, we can still t…
- d608dfa #51438 — [Bugfix][MRV2] Reserve spec-decode lookahead blocks in V2 warmup (#51438)
- 作者: Nick Hill | +449/-37 | 4 个文件
This is a copy of https://github.com/vllm-project/vllm/pull/50531 from @rchalamala.
- 1c1077c #49876 — [Bugfix][Parser] Confirm reasoning end when an Inkling content block opens (#49876)
- 作者: Vegetog | +57/-11 | 2 个文件
Fixes the streaming text-then-tool half of #49865. Rebased onto current main (including #51391) and narrowed to that one bug, per @bbrowning’s review. Inkling’s reasoning pass leaves its reasoning phase only on an explicit reasoning-end event. A response that opens with visible content and no thinking block never emitted one, so DelegatingParser never entered its tool phase and the tool pass never…
- 75231ef #51391 — [Bugfix][Parser] Prevent Inkling block-end leakage with tools (#51391)
- 作者: cherry77-cloud | +352/-30 | 3 个文件
Fixes #51387. ## Problem When tools are enabled, an Inkling plain-text response can leak <|end_message|> or <|content_model_end_sampling|> into content. This commonly occurs on the synthesis turn after a tool result, when the model answers without opening another tool block. Both streaming and non-streaming paths are affected. ## Root cause The reasoning pass runs with tool parsing disabled and fo…
- 58d3918 #51468 — [BugFix] Preserve divergent FA hits with external Mamba state (#51468)
- 作者: Jeff (Junze) Ma | +90/-12 | 3 个文件
The original fix is by @ywang96, with a follow-up refinement by @ivanium truncate_computed_blocks() asserts that every KV cache group holds at least num_computed_tokens // block_size blocks. For a hybrid (full-attention + Mamba) model served with a KV connector, a Mamba group can legitimately hold fewer: its blocks are whole recurrent state, not per-token KV, and an external hit can supply the sta…
- 0fe9a91 #51495 — [Bugfix] Fix LFM2 ShortConv prefix breaking quant ignore list (#51495)
- 作者: ylsun | +2/-2 | 2 个文件
ShortConv is stored on the attribute short_conv but was constructed with prefix=f"{prefix}.conv". The quantization config’s ignore list is rewritten through hf_to_vllm_mapper (".conv." -> “.short_conv.”), so ignored conv projections never match the layer prefix and get built with a quantized linear method, failing to load unquantized weights. ShortConv is stored on the attribute short_conv in the …
⚡ Performance
- 9b0afeb #51458 — [Perf] Avoid some more unnecessary GPU<->CPU syncs (#51458)
- 作者: Nick Hill | +99/-49 | 13 个文件
Split out from https://github.com/vllm-project/vllm/pull/43107. Each of these blocks the calling thread on a path that runs per forward pass. Found by running CI with VLLM_GPU_SYNC_CHECK=error; the check itself and the places where a sync is deliberate are not included here. Replace blocking torch.tensor(…, device=
) construction, which copies from pageable host memory, with async_tensor_h2d… - 9e6be4a #50365 — [Perf][Sparse MLA] Drop the atomic contention in the index remap (#50365)
- 作者: Nick Hill | +87/-33 | 3 个文件
The sparse-MLA index remap splits each 2048-wide row into 16 column tiles. On the valid-count path every tile atomic-adds into the same per-row counter, so the row total is built under 16-way contention – and the DCP compaction path uses that same counter as an atomic slot allocator. Counting is the only reason the tiles need to talk to each other. Give one program the whole row instead: the coun…