共 62 个 commit,涉及 192 个文件,+8110/-1430 行变动。

概要

统计项数值
Commit 数62
变更文件192
新增行数+8110
删除行数-1430

Commit 列表

🖥️ Kernel

  • 66b3c0e #50294 — [Kernel][Model] Optimize FA4 mm_prefix range lookup (#50294)
    • 作者: Li Juqi | +819/-45 | 5 个文件

    [Kernel][Model] Gemma4: optimize FA4 mm_prefix range lookup and CuTe JIT stability Gemma4 multimodal models can use FA4 as the backend for all layers, so that both sliding-attention layers (head_dim=256) and global/full layers (global_head_dim=512) use one attention backend. This is required for the vision mm_prefix / PrefixLM bidirectional mask path, and avoids mixing FlashAttention and Triton …

  • c25093a #40372 — [Kernel] Batch invariant NVFP4 MoE using cutlass (#40372)
    • 作者: Jakub Zakrzewski | +316/-102 | 5 个文件

    Alternative to #39520 . Much like in case of nvfp4 linear, the current VLLM_CUTLASS MoE backend happens to be already batch invariant in practice. This PR adds a few comments and static asserts in the .cu file so that no one breaks it accidentally, along with a specific test case for batch invariance. pytest tests/v1/determinism/test_nvfp4_batch_invariant_cutlass_moe.py ## Test Result ~Note: the t…

  • 3756bf1 #49792 — [Kernel][SM100] Add a CuTeDSL fused query kernel (#49792)
    • 作者: xiaozhoupy | +699/-1 | 6 个文件

    What this does Add an SM100 CuTeDSL implementation of the fused query preprocessing used by DSA sparse attention (fused_q): the packed fp8 MQA query, the indexer query RoPE + UE8M0 fp8 quant, and the folded index weights — registered as a torch custom op. Selected only for the supported dtype/shape combination (is_fused_q_cutedsl_supported); the existing Triton implementation stays the fallback…

  • edbc496 #50697 — [Kernel][Inkling] Fuse shared-expert partial addition into the Lamport collective (#50697)
    • 作者: Canlin Guo | +48/-6 | 3 个文件

    Before: After: - Benchmark for kernel only(The code below is generated by AI for testing) - E2E ## Test Result Kernel: E2E: This PR: Main: —

  • 52c0e3c #50157 — [Kernel] Add support for Flashinfer Mamba SSU algorithm selection (#50157)
    • 作者: amitz-nv | +88/-5 | 4 个文件

    Allow choosing the Flashinfer Mamba SSU algorithm. The new command line argument is –mamba-ssu-algorithm. On some use-cases of Nemotron 3 Nano NVFP4 (like ISL/OSL=1k/8k), using “horizontal” instead of “auto” (which eventually chooses “vertical”) improves throughput. On Nemotron 3 Nano NVFP4 with & without –mamba-ssu-algorithm horizontal: * Throughput benchmark with aiperf (3 runs to account for …

📦 Other

  • 62c5e21 #50507 — [KV Offloading] Support partial-tail prefix reuse with fine-grained prefix matching (#50507)
    • 作者: Chauncey | +360/-44 | 2 个文件

    see https://github.com/vllm-project/vllm/issues/45702 In hybrid Attention–Mamba models such as Qwen3.6, physical KV blocks are often very large to accommodate the Mamba state. Even with a smaller prefix_match_unit, native offloading previously could only store and restore complete physical blocks. As a result, reusable prefix tokens already computed near the end of a block were lost. This change e…

  • 397094d #49990 — Resolve revision to commit_hash once per model load, via huggingface_hub’s resolve_revision (#49990)
    • 作者: Lucain | +84/-6 | 9 个文件

    TL;DR: resolve revision and remote code_revision commit hashes only once via huggingface_hub’s new resolve_revision, preventing repeated downstream resolution and avoiding concurrency issues if a repository is updated while a model is loading Disclaimer: AI-assisted PR, heavily human-reviewed/amended. I’m maintainer of the huggingface_hub which handles all the download/cache system. cc @hme…

  • 5bf3260 #50448 — [Rust Frontend] Deduplicate request preprocessing for /tokenize (#50448)
    • 作者: Sage | +170/-50 | 5 个文件

    /tokenize currently repeats prompt handling already used by inference. This can make its token IDs drift from what generation receives, especially after chat rendering or multimodal prompt expansion. This pr expose and reuse the chat and text request processors from vllm-project/vllm#49045.

  • fd859d2 #44956 — [KV Connector][Mooncake] Add store group semantics (#44956)
    • 作者: Schatten | +531/-3 | 2 个文件

    Mooncake added optional grouped object semantics in kvcache-ai/Mooncake#2180, motivated by the group-semantics discussion in kvcache-ai/Mooncake#2127. vLLM’s MooncakeStoreConnector currently writes each physical Mooncake object independently, even when multiple physical objects belong to the same logical vLLM KV cache entry. This PR adds an opt-in vLLM-side integration so related Mooncake objects …

  • 999dd8b #50526 — [XPU] Alias is_current_stream_capturing to XPU in cuda wrapper (#50526)
    • 作者: Sundaresan G | +1/-0 | 1 个文件

    The XPU backend aliases torch.cuda.* APIs to their torch.xpu.* equivalents in _torch_cuda_wrapper so that CUDA-graph-capturing code paths work on XPU. The graph-related aliases (torch.cuda.graph, CUDAGraph, graph_pool_handle) were already set, but torch.cuda.is_current_stream_capturing was missing, so any code that guards on stream capture hit an AttributeError on XPU. This adds the missing alias …

  • 41a7e7d #50321 — [KV Offload] Support partial secondary-tier load results (#50321)
    • 作者: Moein Khazraee | +88/-6 | 3 个文件

    This makes secondary-tier fetches best-effort and less tightly coupled to an earlier availability snapshot: losing a few keys only requires recomputing those KV values, while every requested value available at fetch time can still be reused. The surrounding design already handles lookup, primary-tier allocation, and completion per key. Keys are combined into asynchronous jobs only for transfer eff…

  • 6153dbe #50940 — [R3] Unify routed expert shape configuration (#50940)
    • 作者: aoshen02 | +146/-46 | 7 个文件
    • preserve and include the Gemma 4 fix from #50460 as the first commit, with yihengz as its original author - normalize experts-per-token aliases into ModelArchitectureConfig.num_experts_per_token - make both the R3 worker capturer and scheduler manager use canonical layer, expert, and top-k counts - fix Kimi K3 R3 initialization through its num_experts_per_token field - reuse the same accessor in…
  • 0187f4c #50841 — [CPU] Enable tcmalloc for s390x (#50841)
    • 作者: Rehan Khan | +39/-2 | 4 个文件

    Enable tcmalloc for s390x Run vllm server and run inference ## Test Result —

  • c416f15 #50404 — [Model] Fix Kimi-K3 MLA with disabled context parallelism (#50404)
    • 作者: Dr Tasos Varoudis | +5/-0 | 1 个文件

    Fix Kimi-K3 MLA decoding when context parallelism is disabled. The shared MLA implementation initializes dcp_world_size with the internal disabled sentinel -1. The regular MLAAttention wrapper resolves that value before invoking the backend, but Kimi-K3 constructs and calls the backend implementation directly. As a result, the Kimi-K3 path can pass invalid context-parallel metadata to the FlashAtt…

  • 9cd9f8c #51068 — Prune redundant tests points in correctness_e2e/[test_sequence_parallel,test_async_tp] (#51068)
    • 作者: Michael Goin | +113/-199 | 5 个文件

    Sequence Parallel Correctness Tests takes over an hour and isn’t a default feature users run into. Let’s greatly simplify the duplicated test points and run it on the nightly. Similar to AsyncTP ## Test Result —

  • 33c5058 #50806 — [ROCm] Restore Inkling MTP backend parity (#50806)
    • 作者: Andreas Karatzas | +42/-13 | 1 个文件
    • Synchronize the backend-neutral AMD MTP orchestration with the NVIDIA implementation updated by PR #48892. - Select and construct only the checkpoint depth layers needed by the configured speculative-token count. - Route each speculative step to its matching depth and ignore checkpoint weights for depths that were not constructed. - Preserve AMD-specific kernels through the existing relative imp…
  • d0ce3da #50607 — [ROCm]: Bump torch 2.12, triton 3.7, torchaudio, torchvision (#50607)
    • 作者: Rohan Potdar | +21/-14 | 2 个文件

    Test Result —

  • 05b7876 #50510 — [MoE][Humming] Support SiTU activation for Kimi-K3 (#50510)
    • 作者: Julian Huang | +1/-0 | 1 个文件

    Enable the Humming expert backend for Kimi-K3’s SiTU activation. ## Test Result —

  • 8ae8337 #50911 — [Spec Decode] Enable fused non-causal TokenSpeed MLA for DSpark (#50911)
    • 作者: NVShreyas | +32/-7 | 2 个文件

    Enable the TokenSpeed MLA backend for DSpark’s non-causal multi-token draft blocks. DSpark produces its draft positions in one non-causal forward pass, but TokenSpeed did not advertise that capability and vLLM did not forward the metadata’s causal mode to the kernel. This prevented the draft model from using TokenSpeed’s fused multi-token path. This PR: - advertises non-causal and non-causal multi…

  • d31de3c #50912 — [Kimi K3 Perf] option to shard the shared expert for non mega case, 16.98 GiB memory/GPU saved (#50912)
    • 作者: Wentao Ye | +10/-18 | 3 个文件

    A following up PR for https://github.com/vllm-project/vllm/pull/50656 @tlrmchlsmth config | effectiveness – | – VLLM_KIMI_K3_SHARD_SP_SHARED_EXPERT=0 | No non-mega、TP=4、DP=1 | No non-mega、TP=4、DP=2、EP | Yes MegaMoE、TP>1、EP | Yes PP>1 | No The update is safe as we make the shared experts overlap off, this might be a future optimization point Recommend to use with PD disaggregation (decode node) M…

  • 7cab436 #50929 — [MM][CG] Support ViT full CUDA graph for Kimi-K2.5 (#50929)
    • 作者: Linkun | +247/-1 | 1 个文件

    Add ViT CG support (#38175 ) for Kimi K2.5/K2.6 ## Test Result “The user wants me to describe the image. Let me look at the image carefully. The image shows a soccer/football” —

  • 7635a90 #49919 — [Core] Explicitly manage torch CPU threads in workers (#49919)
    • 作者: Nick Hill | +207/-25 | 4 个文件

    Clamp for startup, single thread for serving. torch defaults its intra-op thread pool to the host core count, ignoring process affinity, cgroup CPU quotas, and co-located worker processes. Any torch CPU op above the parallel_for grain size (32k elements) fans out across that pool, and the OpenMP workers then spin-wait after every parallel region - stealing cycles from the engine loop’s serial code…

  • 7ac2ec7 #50593 — [Kimi-K3][AMD] Fuse AttnRes state updates and norms (#50593)
    • 作者: yinfengLiu | +208/-49 | 3 个文件

    The AMD Kimi-K3 decoder previously executed several memory-bound operations around AttnRes as separate kernels: 1. Update the accumulated prefix residual. 2. Write the updated prefix into block-residual storage. 3. Run the AttnRes weighted residual aggregation. 4. Apply the input RMSNorm. 5. Add the self-attention output to the prefix. 6. Run the MLP AttnRes aggregation. 7. Apply the post-attentio…

  • 166f4e2 #50867 — fix: fuse weightless RMSNorms at their declared width (#50867)
    • 作者: Anuj Bolewar | +29/-23 | 2 个文件

    RMSNormFuser.fuse() sizes a weightless norm (one without a weight parameter) from model_config.get_hidden_size(). That is the LM hidden size and is wrong for norms on a sub-dimension — the exact mismatch in #39061, where a 72-wide vision-encoder norm was fused at 5376. The fused norm’s weight buffer is only used to size the norm and to drive the TP-shard gather in TPAwareNormMixin. With the wrong …

  • 4f819f8 #38390 — [Model Runner v2] E/P/D disaggregation support (#38390)
    • 作者: Wentao Ye | +145/-8 | 7 个文件

    E/P/D disaggregation support for MRv2 ### Single image

  • f9c74b4 #50991 — [Mamba] enable prefix cache by default (#50991)
    • 作者: Jiangyun Zhu | +30/-36 | 4 个文件

    enable prefix cache for hybrid model by default and change the default mode to align ## Test Result —

  • 5789897 #49969 — [Spec Decode] Add top-k DSpark Markov projection (#49969)
    • 作者: Andrii Skliar | +235/-26 | 4 个文件

    DSpark computes the base logits for all draft positions in parallel, then applies the Markov bias sequentially over the full draft vocabulary. The repeated full-vocabulary Markov projection is on the serial drafting path. This PR adds dspark_draft_topk for Qwen3 DSpark. It selects the top-k candidates from the base logits once, gathers the corresponding markov_w2 rows, and evaluates the sequential…

  • 1eb3694 #51014 — [Docs] Fix two docs build warnings (#51014)
    • 作者: Harry Mellor | +2/-2 | 2 个文件

    Fixes two warnings emitted by mkdocs build: 1. Broken cross-reference in docs/design/moe_kernel_features.md. The naive row of the all2all backend table pointed at vllm.model_executor.layers.fused_moe.runner.MoERunner, but runner/init.py re-exports nothing — the class is defined in runner/moe_runner.py. The link rendered as plain text for readers. The link text (layer.py) was also stale fro…

  • 59b2fdf #48250 — Support MLA properly in the Transformers modeling backend (#48250)
    • 作者: Harry Mellor | +699/-52 | 9 个文件

    Requires the following Transformers side PRs: - https://github.com/huggingface/transformers/pull/47435 - https://github.com/huggingface/transformers/pull/47451 - https://github.com/huggingface/transformers/pull/47460 Portions of this PR that were split into their own PRs: - https://github.com/vllm-project/vllm/pull/49957 - https://github.com/vllm-project/vllm/pull/49982 — vLLM side changes: - Ad…

  • 24c939c #47104 — [XPU] fix collecting oneccl version info (#47104)
    • 作者: Yan Ma | +2/-3 | 1 个文件

    oneccl version info on XPU platform should not be collected from apt lib. Here we refer the mpirun –version as it reflects the correct oneCCL system really uses. python vllm/collect_env.py ## Test Result —

  • e98a877 #50368 — [Rust Frontend][gRPC] Add multimodal image inference (#50368)
    • 作者: Connor Carpenter | +553/-69 | 12 个文件
    • Add image inputs to the Rust frontend gRPC Inference service. - Extend GenerateRequest with media items supporting HTTP(S) URLs, data URIs, and raw bytes. - Reuse the existing multimodal preprocessing pipeline to fetch and process images, expand placeholder tokens, and attach engine-facing multimodal features. - Preserve media UUID and MIME type metadata. - Require token-ID prompts when media is…
  • 0b1c151 #50540 — [Rust Frontend] Align tool rendering for Kimi K3 (#50540)
    • 作者: Bugen Zhao | +85/-9 | 4 个文件

    Align the Rust Kimi K3 renderer’s tool declarations with the Python and checkpoint encoding behavior. Related Python fix: #50228. That PR fixes the corresponding K3 renderer path in the Python frontend; this PR covers the Rust renderer behavior only. The renderer previously serialized a missing function description as an empty string and omitted strict for every tool. This change: - omits descript…

🐛 Bug Fix

  • 08b8613 #48929 — [Bugfix][Model] Fix MiniMax-M3 NVFP4 inference correctness (#48929)
    • 作者: Gabriel Wu | +100/-31 | 2 个文件

    Fix two MiniMax-M3 correctness issues exposed by the NVFP4 checkpoint. First, the routed experts use packed SWIGLUOAI_UNINTERLEAVE with model-specific alpha, beta, and clamp values. FlashInfer CUTLASS already supports this math, but the vLLM adapter neither advertised the packed activation nor forwarded all three parameters. Marlin similarly replaced missing quant-config alpha/beta values with pla…

  • cd930c8 #38771 — [Bugfix] Fix MLA kv_b_proj activation dtype with Marlin FP8 (#38771)
    • 作者: Jacob Zhang | +24/-25 | 1 个文件

    Fixes #38658. This PR fixes an MLA prefill dtype bug when FP8 weights are served through the Marlin path on GPUs without native FP8 support (sm < 89). On affected GPUs, Marlin repacks FP8 weights into torch.int32. In vllm/model_executor/layers/attention/mla_attention.py, compute_prefill_context() was using self.kv_b_proj.weight.dtype to determine how to cast kv_c_normed before passing it to kv_b…

  • a9b39d6 #51153 — [Bugfix] Enable chunked prefill for qwen3.5-0.8B ppl test (#51153)
    • 作者: music-dino | +3/-0 | 1 个文件

    PR #50991 enabled prefix caching by default and changed the default mode to align, and updated several hybrid model tests to enable chunked prefill to go along with the new defaults. It missed models/language/generation_ppl_test/test_qwen.py::test_ppl[model_info2] so the test failed in the Language Models Test (PPL) on the nightly. This PR enables chunked prefill for the failing hybrid model test …

  • beca88e #51131 — [BugFix][K3] Skip moe_intermediate padding when EP is enabled (#51131)
    • 作者: Ziming Huang | +1/-1 | 1 个文件

    FIX https://github.com/vllm-project/vllm/issues/51124 ## Test Result —

  • f5cd862 #50649 — [ROCm][Bugfix] Kimi-K3 Fix KDA NaN on mixed batches and racy autotune config (#50649)
    • 作者: kliuae | +556/-7 | 3 个文件

    For Kimi-K3, there are two correctness bugs on the ROCm Kimi-K3 KDA path. - KimiGatedDeltaNetAttention._forward sends the decode sequences of a mixed decode+prefill non-speculative batch through the KDA chunk kernel chunk_kda_with_fused_gate. The chunk kernel treats each decode as length=1 sequence, and returns NaN for them. This PR separates the decode out of the mixed batch and pass them to fuse…

  • 4719a9b #51092 — [Bugfix][Spec Decode] Fix EAGLE3 DeepSeek draft crash on non-YaRN rope configs (#51092)
    • 作者: qizixi | +0/-9 | 1 个文件

    DeepseekV2Eagle3DecoderLayer.init rebuilds config.rope_parameters from config.rope_scaling before handing the config to DeepseekV2MLAAttention: DeepseekV2MLAAttention already performs exactly this normalization itself (deepseek_v2.py), so for a YaRN draft the block is a no-op. But it is not harmless: 1. It crashes plain-RoPE EAGLE3 drafts. Under Transformers v5 config.rope_scaling is a dep…

  • 50c5168 #50275 — [Bugfix][EC Connector] Don’t stop an encoder-instance request before its images are encoded (#50275)
    • 作者: Tianyu Guo | +121/-5 | 4 个文件

    Problem On an EPD encoder instance a request can be marked finished before its multi-modal items are encoded. Under sync scheduling this is silent: no embedding reaches the EC connector, the decode instance re-encodes locally, and the client still gets a successful completion. Under async scheduling it kills the engine. The fabricated token decrements num_output_placeholders for a reque…

  • 7794b1e #49397 — [Bugfix] Skip Qwen3 deepstack buffers without vision (#49397)
    • 作者: Wei-Cheng (Wayne) Chiu | +64/-61 | 3 个文件
    • avoid creating Qwen3 deepstack tensors when the vision tower is skipped - preserve zero-backed deepstack inputs when vision is enabled, keeping the compiled decoder input contract stable - apply the same fix to Qwen3 Omni, Qwen3-VL, and Qwen3-VL-MoE Fixes #49384. ## Why this is not duplicate work I searched open PR titles and bodies for #49384, the reported Qwen3-Omni meta failure, StageMissingL…
  • eb3dce9 #50958 — [Bugfix][Model] Gemma3n/Gemma4: pad variable-length audio batches (#50958)
    • 作者: TrainToGPB | +165/-11 | 4 个文件

    Fixes #50957. Four concurrent requests, one audio clip each, kill EngineCore when the clips differ in duration: Every in-flight request then fails with EngineDeadError, so one mixed-duration batch takes the server down. _call_hf_processor unpads audio features per item so a multimodal cache entry does not depend on the batch it was first processed in, and its comment claims _get_mm_fields_config r…

  • 2cb3ff8 #50405 — [BUGFIX][Quant]Fix test_kv_scale_reload failed (#50405)
    • 作者: Yejing Lai | +3/-2 | 1 个文件

    Fix UT tests/model_executor/model_loader/test_reload.py::test_kv_scale_reload failed. error msg: The size of tensor a (16384) must match the size of tensor b (2) at non-singleton dimension 1 Root cause: PR #41652 set self.strategy=CHANNEL(see compressed_tensors_w8a16_fp8.py#L140), it will cause the per-tensor scale not to be replaced with the expected channelwise output when reloading the checkpoi…

  • 77b5199 #51050 — [CI][Bugfix] Fix test_shutdown_on_engine_failure startup deadlock (#51050)
    • 作者: Nick Hill | +39/-39 | 1 个文件

    The test polled for server readiness without ever draining the server’s stdout/stderr pipes, so a server whose startup output exceeded the pipe buffer blocked on write and never became ready, failing with “Server failed to start in 120 seconds”. Redirect the server’s output to a file instead. This also removes the ROCm special case, which disabled pipe capture to dodge the same hang, and lets both…

  • ffee324 #48061 — [BugFix][Mooncake] Use global data_parallel_index for the DP engine index (#48061)
    • 作者: Yifan Qiao | +3/-16 | 3 个文件

    get_mooncake_dp_engine_index asserted data_parallel_rank_local is not None under external/hybrid DP LB. This crashes on headless wide-TP follower nodes: they go through run_headless → MultiprocExecutor directly, which never populates data_parallel_rank_local. Preferring data_parallel_rank_local was also inconsistent: for a replica launched with –data-parallel-rank N (N > 0), head-node processes c…

  • 7153fd7 #49558 — [Bugfix][MoE] Filter packed expert weights during EP loading (#49558)
    • 作者: aoshen02 | +16/-1 | 2 个文件
    • Treat per-expert .weight_packed tensors as heavy weights eligible for EP-local filtering. - Keep scale and metadata tensors from every expert unchanged. - Cover local packed weights, remote packed weights, and remote scale tensors in the existing unit suite. ## Root cause The EP filter previously recognized only names ending in .weight. Quantized expert checkpoints can store the primary expert p…
  • 122b3d4 #50915 — [Bugfix][CPU] Fix macOS build: std::sqrt is not constexpr under libc++ (#50915)
    • 作者: Harjoth Khara | +2/-2 | 1 个文件

    The nightly macOS Apple Silicon Smoke Test has failed every night since 2026-07-31. Both runners — macos-15 (required) and macos-26 — fail while building: Latest failing run: https://github.com/vllm-project/vllm/actions/runs/30782456002 ### Why it happens The code asks for a compile-time constant built from std::sqrt: std::sqrt is not constexpr until C++26. GCC allows it anyway as an extension…

  • 199644d #50906 — [Bugfix][Attention] Guard sparse MLA masked MHA workspace (#50906)
    • 作者: yimdev | +115/-13 | 3 个文件

    Sparse MLA masked-MHA builds a bit-packed top-k mask whose size is approximately: The existing workspace was fixed at 64 MiB, while the routing matrix allows a single 32K prefill to enter masked-MHA. Such a request requires exactly 128 MiB. Heterogeneous batches can require even more despite staying within max_num_batched_tokens. Attempting to reshape the fixed workspace for an oversized mask can …

  • 7c40d61 #50250 — [Bugfix] Flatten >2D multimodal embeddings, not just 3D (#50250)
    • 作者: Michał Ganczarenko | +2/-2 | 1 个文件

    MultiModalMixin._split_embeddings (Transformers modeling backend, vllm/model_executor/models/transformers/multimodal.py) only flattened vision/audio encoder output to 2D when embeddings.ndim == 3. CohereLabs/aya-vision-8b’s get_image_features returns embeddings shaped [num_items, H, W, hidden] (4D — the spatial grid is never flattened by the model), so the tensor skipped the flatten step and reach…

  • 5f2ee2f #50462 — [Bugfix][Core] Log KV cache capacity after block-size resolution (#50462)
    • 作者: Tanmay Dixit | +19/-23 | 2 个文件

    Fixes #50456. KV-cache capacity was logged during worker KV-cache configuration, while the authoritative values exposed through cache_config.kv_cache_size_tokens and kv_cache_max_concurrency are published later from the resolved scheduler configuration. This PR centralizes those observability updates after scheduler KV-cache configuration generation and effective block-size resolution: - removes t…

  • 6a9fdf0 #50950 — [Bugfix] Resolve seq-cls num_labels from the top-level config for multimodal checkpoints (#50950)
    • 作者: Lazurite | +71/-2 | 2 个文件

    Multimodal sequence-classification checkpoints cannot be served: the score head is built with the wrong number of labels, so weight loading fails with AssertionError: Tried to load weights of size torch.Size([20, 2560]) to a parameter of size torch.Size([2, 2560]) Reported against vLLM in #47956 (Qwen3-VL, 20 labels) and downstream as modelscope/ms-swift#9704. ## Root cause as_seq_cls_model.init…

  • a1657a0 #51003 — [Bugfix][Build] Fix DeepGEMM CUDA 12.9 FP8 header visibility (#51003)
    • 作者: Kevin H. Luu | +2/-2 | 2 个文件

    Fix the CUDA 12.9 release build failure in DeepGEMM’s host extension while preserving the exact DeepGEMM f5a76426 code baseline. The pinned mqa_logits.cuh directly names __nv_fp8_e4m3, but it does not include the CUDA header that declares the type. CUDA 12.9 host compilation therefore fails with: This updates both synchronized vLLM DeepGEMM pins from f5a76426 to e21c821f. The new commit has f5a764…

🧪 CI/Tests

  • a3b8675 #51095 — [CI] Fix CI authorization notification fallback (#51095)
    • 作者: Kevin H. Luu | +97/-13 | 5 个文件
    • resolve approval notifications from the source workflow’s head commit when workflow_run.pull_requests is empty or unrelated - let a ready label recover a missed approval notification while retaining marker-based deduplication - give approval-triggered notification runs a unique concurrency fallback when no PR number is attached - run add_label_automerge, notify-ci-authorized, and record-ci-appro…
  • 96e333e #51127 — [CI] Run control-plane workflows on vLLM runners (#51127)
    • 作者: Kevin H. Luu | +2/-2 | 2 个文件
    • run the pre-commit pre-run-check gate on the vllm-runners self-hosted runner group - run the PR-comment CI broker on the same vllm-runners group - keep the existing self-hosted pre-commit runner labels unchanged Both jobs now use the repository’s established selector: ## Why The organization was at its 20/20 standard GitHub-hosted concurrency limit. For the pre-commit workflow, waiting on the Gi…
  • bb543ef #50905 — [ROCm][CI] Add aiter per-token FP8 quant roundtrip and RMSNorm determinism tests (#50905)
    • 作者: Divakar Verma | +31/-8 | 1 个文件
    • Add test_per_token_quant_matches_native: verifies per-token FP8 quantization roundtrip accuracy (per-tensor version existed, per-token did not) - Add test_rms_norm_determinism: verifies AITER RMSNorm produces bitwise-identical results across repeated calls
  • d16e7f9 #51069 — [CI] Prune PyTorch Compilation Unit Tests (#51069)
    • 作者: Michael Goin | +124/-172 | 3 个文件

    PyTorch Compilation Unit Tests takes way too long. It currently has timeout_in_minutes: 150 ## Test Result —

  • 43c4bdc #51087 — [CI] Add run-all comment commands (#51087)
    • 作者: Kevin H. Luu | +72/-14 | 3 个文件

    Add two exact, intentionally undocumented CI comment variants for trusted users and authorized PR authors: - /ci run all triggers a Buildkite build with RUN_ALL=1. - /ci run nightly triggers a Buildkite build with RUN_ALL=1 and NIGHTLY=1. Both variants use the existing authorization, deduplication, PR-head validation, reactions, and acknowledgement path. The GitHub Actions workflow continues to ma…

  • 12292d9 #50323 — [CI] Detect and fail evals on when NaNs appear in logits (#50323)
    • 作者: Tyler Michael Smith | +60/-3 | 6 个文件

    This PR sets VLLM_COMPUTE_NANS_IN_LOGITS=1 in several evals so we can test against NaNs. NaNs appearing in logits often co-appear with KV cache NaNs, which is a catastrophic failure mode for a inference service since attention kernels mask with multiplication by zero. (see https://github.com/Dao-AILab/flash-attention/issues/1974) Assisted-by: OpenAI Codex

  • c687c1a #51079 — [ci] Update CI notify workflow with PR write permissions (#51079)
    • 作者: Kevin H. Luu | +1/-1 | 1 个文件
  • a5149b2 #51015 — [CI] Stabilize GLM-5.2 PCP evaluation (#51015)
    • 作者: Kevin H. Luu | +1/-0 | 1 个文件
    • enable PyTorch expandable CUDA allocator segments for the GLM-5.2 TP1/PCP4 GSM8K evaluation; - keep –max-num-batched-tokens 32768, which is needed to exercise the large-batch PCP path; - scope the allocator change to the memory-tight TP1/PCP4 case; TP2/PCP2 is unchanged. ## Root cause The pytest-level error in Buildkite #81913 is only a 20-minute server-start timeout. The worker log shows the a…

⚡ Performance

  • b92352c #50992 — [Perf][KV Offload] Avoid quadratic ARC batch eviction (#50992)
    • 作者: MINJUN GIL | +132/-21 | 2 个文件

    ARC batch eviction repeatedly scans its internal cache lists from the beginning for each block selected. As a result, evicting many blocks can require quadratic work. ## Fix Keep monotonic iterators over those lists while collecting candidates, so each entry is visited at most once. Cache mutations remain deferred until all requested candidates are found, preserving atomic eviction. This also pres…

🦀 Rust Frontend

  • 7743486 #50580 — [Frontend] DeepSeek V4 0731 reasoning effort prompts & mappings (#50580)
    • 作者: Bugen Zhao | +205/-51 | 5 个文件

    Align DeepSeek V4 Flash prompt rendering with DeepSeek-V4-Flash-0731 and the current hosted API’s model-specific effort mapping. The renderer accepts the canonical 0731 levels: low emits no prefix, high emits the “Absolute maximum” prefix, and max emits the “Beyond maximum” prefix. DSML parsing and output syntax retain their existing behavior. The request normalization follows the public deepseek-…

🔩 Misc

  • 7b50d2c #50879 — [Misc] Avoid importing nixl_ep on every vllm serve config (#50879)
    • 作者: Nicolò Lucchesi | +24/-23 | 1 个文件

    Avoid importing nixl_ep all the time we boot up vllm. Instead, only do it lazily when needed (dpep configuration). Right now if you do vllm serve … import_utils will try to resolve nixl_ep, and in this case even log “unrelated” import failures as this configuration can then go on and run just fine, as no DPEP is needed (I have no libcudart 12 on this machine, but it only matters if it’s actual…