共 20 个 commit,涉及 95 个文件,+2817/-440 行变动。
概要
| 统计项 | 数值 |
|---|---|
| Commit 数 | 20 |
| 变更文件 | 95 |
| 新增行数 | +2817 |
| 删除行数 | -440 |
Commit 列表
⚡ Performance
- 83f591d #51967 — [Perf][DSV4] Optimize global top-k index kernel with compile-time constants (#51967)
- 作者: Chauncey | +5/-5 | 1 个文件
Optimize global top-k index kernel with compile-time constants ## Test Result ## Serving Benchmark Results | Version | Mean Output Throughput | Mean TPOT | Completed Requests | Relative Change | | — | —: | —: | —: | —: | | main | 634.97 tokens/s | 103.33 ms | 128/128 | Baseline | | this pr | 638.12 tokens/s | 102.32 ms | 128/128 | Throughput +0.50%, TPOT -0.98% | ## Kernel Microb…
- 836aac9 #52084 — [Perf][DSV4] Optimize sparse top-k metadata kernels for higher prefill throughput (#52084)
- 作者: Chauncey | +1/-1 | 1 个文件
Optimize sparse top-k metadata kernels for higher prefill throughput | Tokens | 128 Workers | 256 Workers | Relative Change | | —: | —: | —: | —: | | 1,024 | 9.75 us | 8.27 us | 15.2% faster | | 4,096 | 22.40 us | 15.52 us | 30.7% faster | | 16,384 | 74.76 us | 47.14 us | 36.9% faster | ### Paired A/B/A Serving Results | Run | Workers | Mean Output Throughput | Mean TPOT | | –…
📦 Other
- 9409f59 #52514 — [Core] Add CuMemAllocator.discard() for tag-selective GPU memory release (#52514)
- 作者: Dakai An | +132/-0 | 4 个文件
Follow-up to https://github.com/vllm-project/vllm/pull/46438. PR #46438 introduces tag-selective discard(), allowing stale allocations such as KV cache to be released while keeping model weights mapped and usable. Based on 46438, This PR refines the interaction between sleep() and discard() for both CUDA and XPU allocators: - Keep mapped weights usable when only KV-cache tags are discarded. - Sync…
- fa9d67f #49585 — [EC Connector] Added Build Connector Worker Meta for EC Connector (#49585)
- 作者: omerpaz95 | +430/-16 | 16 个文件
Why this is still needed after #38390. The PR implemented the V2 model runner EC Connector, but EC still has no worker -> scheduler metadata channel. KV connectors have a complete one: build_connector_worker_meta() -> KVConnectorOutput.kv_connector_worker_meta -> KVOutputAggregator -> scheduler-side KV connector. EC had none of those three pieces, so a worker-sideECConnector has no way to repo…
- d480199 #51901 — [CI/Build] Add warning for unsupported global PTX architecture requests in… (#51901)
- 作者: Shane Widanagama | +63/-5 | 4 个文件
… CMake configuration. Implements one item from #9129: Warn that PTX builds are not currently supported (post [CI/Build] Per file CUDA Archs (improve wheel size and dev build times) #8845), currently if there is a +PTX in TORCH_CUDA_ARCH_LIST this will be ignored. We should warn when this is the case Users can request PTX through TORCH_CUDA_ARCH_LIST values such as 8.0+PTX. vLLM strips the Torch…
- 4215646 #52384 — [Rust Frontend][gRPC] Preserve skip_special_tokens decoding option (#52384)
- 作者: Biswa Panda | +5/-1 | 2 个文件
Align the native Generate gRPC API with vLLM’s existing Python serving APIs by carrying the per-request skip_special_tokens decoding option through ResponseOptions. The Python chat completions, completions, responses, and token-native APIs default this option to true, but the gRPC request did not expose it. Consequently, a gRPC caller could not explicitly preserve tokenizer-defined special markers…
🐛 Bug Fix
- fe1c317 #52356 — [Bugfix][ROCm] Skip FP8 MLA prefill PS-metadata build for chunked-context batches (#52356)
- 作者: Shantipriya Parida | +5/-1 | 1 个文件
FP8 MLA PS ASM kernel is run only on non-chunked (causal) prefills. However, its persistent metadata is always built, even during chunked prefills (unnecessarily). This PR fixes it so metadata is only computed when it’s actually needed. Unit test for the exact guarded call site: Formatting/lint: End-to-end serving benchmark comparing an unpatched build against this PR applied on top of the identic…
- 4d2a68d #52436 — [Bugfix][Spec Decode][Structured Output] DSpark: fix the grammar bitmask mapping when the draft budget is zero (#52436)
- 作者: oops-oom | +89/-15 | 3 个文件
Fixes a crash when DSpark adaptive verification (enable_adaptive_verification, added in #47808) is combined with structured outputs and the chosen draft budget is zero. Scenario: serving a structured-outputs request (JSON schema / grammar) with DSpark adaptive verification enabled crashes on an assert as soon as the drafter’s confidence drops far enough that adaptive verification decides to ve…
- 84530eb #52441 — [Bugfix][Multimodal] Keep Gemma 4 video frame counts on CPU (#52441)
- 作者: Chauncey | +7/-1 | 2 个文件
FIX https://buildkite.com/vllm/ci/builds/84014#01a0041b-4750-418e-aab2-15f6f61ebfaf ## Test Result —
- 1b079c4 #52311 — [Bugfix][Model Runner V2][Spec Decode] Fix off-by-one in bad_words draft-prefix matching (#52311)
- 作者: Jyan-R | +108/-1 | 2 个文件
Fix an off-by-one in the spec-decode branch of bad_words_kernel (vllm/v1/worker/gpu/sample/bad_words.py), present since the kernel was introduced in #33433. The sampler passes input_ids gathered at logits_indices, so for a spec-decode request the per-request local layout is: local position 0 = last committed sampled token, local position j = draft token d{j-1}. The sibling kernels index draft to…
- 70aaec8 #52246 — [Bugfix][Anthropic] Return 4xx for client-caused errors in /v1/messages (#52246)
- 作者: Do_it_now_! | +100/-20 | 3 个文件
Fixes #52088 Invalid user input that passes the Anthropic request schema but fails the Anthropic→OpenAI conversion (e.g. more than VLLM_MAX_STOP_STRINGS stop sequences) currently surfaces as an HTTP 500: the except Exception in vllm/entrypoints/anthropic/api_router.py swallows the pydantic.ValidationError raised while constructing ChatCompletionRequest and turns it into an internal_error. The sche…
- 41f12a0 #52394 — [Bugfix] Raise
VLLMValidationErrorfrom structured output validators (#52394)- 作者: Jeffrey Wang | +109/-30 | 8 个文件
Structured output validators raise raw ValueError. AsyncLLM.generate re-raises only VLLMClientError untouched and wraps the rest in EngineGenerateError, so a bad response_format schema returns 500 instead of 400. This issue is surfaced from ray serve LLM because vLLM 0.27 wraps validator ValueErrors in EngineGenerateError, which is treated as a server error. Ray then interprets an invalid response…
- 8efa13b #52401 — [Bugfix] Pick the DeepSeek V4 eager cudagraph region per model runner (#52401)
- 作者: Nick Hill | +96/-66 | 3 个文件
#51430 narrowed the DeepSeek V4 eager cudagraph region, which corrupts MRV1 output, and #51768 responded by defaulting the model to MRV2 and rejecting MRV1 + PIECEWISE. That default costs ROCm, where MRV1 is still the faster runner for this model. Choose the region from the runner instead: MRV1 wraps the whole attention body in _prepare_and_attn_eager, restoring the pre-#51430 region it needs, whi…
- 6593754 #52419 — [Bugfix][Spec Decode] Keep EAGLE cache registration on the partial-hash-hit path (#52419)
- 作者: mispa-ms | +92/-20 | 2 个文件
HybridKVCacheCoordinator.cache_blocks decides once how far a request may be registered in the prefix-cache hash map. With fine-grained partial hash hits that bound is the raw token count, because a hit no longer has to land on a scheduler_block_size boundary. #50062 rewrote the EAGLE branch to re-derive its own bound from num_finalized_computed_tokens with an unconditional so the rounding comes ba…
- edd4c81 #51318 — [Bugfix][DSv4] Revert adaptive C128A metadata packing (#51318)
- 作者: Toby Mao | +7/-56 | 2 个文件
Revert the adaptive C128A top-k width and packed persistent-buffer views introduced by #50004. The C128A metadata builder runs before FULL CUDA graph replay, but the sparse decode consumer is captured and reuses its capture-time row layout. With #50004, runtime metadata writes packed rows using a batch-dependent active_topk_width, while the captured consumer retained the capture-time row stride. R…
- c94cdd0 #49613 — [Bugfix][Sampling] Clear empty side on thinking-budget asymmetric SWAP (#49613)
- 作者: Henry Su | +97/-2 | 2 个文件
Fix a production bug in ThinkingBudgetStateHolder.sync_batch where swapping a budgeted request with an unbudgeted batch slot leaves stale thinking-budget state at the empty index. When the scheduler/input-batch reorders requests (MoveDirectionality.SWAP), a budgeted request at index i swapped with an unbudgeted request at index j previously: 1. Copied state from i → j via dict.get() + assign 2. Le…
- ed0f475 #52445 — [Bugfix][Model] Kimi-K3 MegaMoE: pass situ_beta/situ_linear_beta to fp8_fp4_mega_moe (#52445)
- 作者: Uranus | +2/-2 | 1 个文件
KimiMegaMoEExperts.forward passed activation_beta=/activation_linear_beta= as keyword arguments to deep_gemm.fp8_fp4_mega_moe. Neither name exists in the signature of deepseek-ai/DeepGEMM at the pinned commit 8b1392b (nv_dev tip), which declares them as situ_beta and situ_linear_beta: The first mega-MoE forward therefore raises TypeError: fp8_fp4_mega_moe() got an unexpected keyword argument ‘acti…
- 97388c4 #51538 — [Bugfix] Make DSV4 sparse MLA work end-to-end for plain decode, MTP, and DSpark (#51538)
- 作者: Gabriel Wu | +797/-120 | 20 个文件
DeepSeek-V4-Flash-0731 could not run reliably through the SM120 sparse MLA backend. This fixes the seven defects that blocked it across all three decode modes – plain decode, MTP, and DSpark – verified end-to-end on 8xRTX PRO 6000 Blackwell across in-flight batching and prefill/decode disaggregation. Commits 1-5 unblock DSpark. Commits 6-7 fix a hang that is not DSpark-specific: it strands a…
- 5cecfc0 #52431 — [Bugfix] Fix modelscope usage (#52431)
- 作者: Cyrus Leung | +3/-5 | 2 个文件
- Fix a KeyError when calling modelscope_list_repo_files, because recent versions of ModelScope no longer return the “Type” field. (I encountered this issue when trying to load meta-models/Muse-Glimmer-30B via ModelScope) - Accept VLLM_USE_MODELSCOPE=1, not just VLLM_USE_MODELSCOPE=True, to be consistent with other boolean-based env vars. ## Test Result —
🦀 Rust Frontend
- ac2ae87 #45802 — [Frontend] Support count_reasoning_tokens in the Streaming Parser Engine (#45802)
- 作者: Chauncey | +669/-73 | 16 个文件
Add token-aware reasoning token counting for the Streaming Parser Engine and surface the count through OpenAI-compatible usage fields. - Adds completion_tokens_details.reasoning_tokens to usage responses. - Propagates token counts through the parser engine pipeline: TokenIDScanner -> IncrementalLexer -> StreamingParserEngine -> SemanticEvent. - Counts only REASONING_CHUNK tokens, excluding reasoni…