共 58 个 commit,涉及 220 个文件,+7523/-2377 行变动。

概要

统计项数值
Commit 数58
变更文件220
新增行数+7523
删除行数-2377

Commit 列表

🐛 Bug Fix

  • 58d3918 #51468 — [BugFix] Preserve divergent FA hits with external Mamba state (#51468)
    • 作者: Jeff (Junze) Ma | +90/-12 | 3 个文件

    The original fix is by @ywang96, with a follow-up refinement by @ivanium truncate_computed_blocks() asserts that every KV cache group holds at least num_computed_tokens // block_size blocks. For a hybrid (full-attention + Mamba) model served with a KV connector, a Mamba group can legitimately hold fewer: its blocks are whole recurrent state, not per-token KV, and an external hit can supply the sta…

  • 0fe9a91 #51495 — [Bugfix] Fix LFM2 ShortConv prefix breaking quant ignore list (#51495)
    • 作者: ylsun | +2/-2 | 2 个文件

    ShortConv is stored on the attribute short_conv but was constructed with prefix=f"{prefix}.conv". The quantization config’s ignore list is rewritten through hf_to_vllm_mapper (".conv." -> “.short_conv.”), so ignored conv projections never match the layer prefix and get built with a quantized linear method, failing to load unquantized weights. ShortConv is stored on the attribute short_conv in the …

  • 6e18901 #51435 — [Bugfix][MM] Avoid device sync in FusedInputNorm initialization (#51435)
    • 作者: Andreas Karatzas | +60/-7 | 2 个文件
    • Determine FusedInputNorm identity from explicit CPU tensors during model construction. - Keep non-identity normalization buffers on the caller’s default device, preserving device-side preprocessing. - Cover identity and non-identity initialization under an accelerator default device. PR #50411 moved multimodal normalization onto the accelerator, which also moved the construction-time torch.allcl…
  • d3621c1 #51076 — [Bugfix][Multimodal] Fix PyNvVideoCodec video backend returning NCHW instead of NHWC (#51076)
    • 作者: Junpu Yu | +39/-1 | 2 个文件

    PyNvVideoCodecVideoBackendMixin._decode_to_pinned_host (vllm/multimodal/video.py) permutes the decoded frame batch to NCHW, but the video-loader contract returns NHWC — as the method’s own docstring states (“frames … in NHWC RGB format … identical to the OpenCV/PyAV backends”) and as the OpenCV/PyAV/TorchCodec/DeepStream backends all do. PyNvVideoCodec RGB DecodedFrames are HWC, so tor…

  • 22a1759 #51469 — [BugFix] Reject invalid data-parallel RPC ports (#51469)
    • 作者: aoshen02 | +24/-8 | 4 个文件

    data_parallel_rpc_port is a fixed bootstrap endpoint shared by independently launched nodes in native MP data parallelism. It defaults to 29550. The previous code treated an explicit value of 0 as an automatic-port request: A producer-to-consumer audit shows two different outcomes. Neither provides a usable port=0 contract. ## Scope: native MP EngineCore startup This path is selected by parallel_c…

  • d21fed9 #51427 — [Bugfix][models_multimodal] Remote HF python code misses importing class (#51427)
    • 作者: Gyula Chinoradszki | +9/-0 | 1 个文件

    Merging of https://github.com/vllm-project/vllm/pull/48413 introduced two independent runtime errors in the Multi-Modal Models (Extended Generation 3) test group (CI 11817).: 1. NameError: name ‘List’ is not defined 2. AttributeError: ‘MiniCPMV’ object has no attribute ‘all_tied_weights_keys’ This PR aims to resolve the first one. The other bug will be fixed in another PR. The root cause w…

  • a828536 #51432 — [Bugfix][multi_modal] Fix pos_ids being unitialized for minicpmv2.6 in hf runner (#51432)
    • 作者: music-dino | +18/-0 | 1 个文件

    #48413 introduced two separate failures in the Multi-Modal Models (Extended Generation 3) TG. - CI The two types of failures are: - NameError: name ‘List’ is not defined - AttributeError: ‘MiniCPMV’ object has no attribute ‘all_tied_weights_keys’ The second one has two parts; this PR addresses the second part of the second failure: - The first issue, presented by the error message is circumvented …

  • 5eaa70d #51411 — [Bugfix][Quantization] Fix INT8 W8A8 MoE crash in TritonExperts (#51411)
    • 作者: djramic | +2/-0 | 1 个文件

    TritonExperts did not recognize torch.int8 as a valid input dtype, causing INT8 W8A8 MoE models to crash when processing quantized activations. The change adds INT8 to the supported activation types and uses BF16 as the compute type, allowing INT8 activations to pass through the Triton MoE path correctly. pytest tests/quantization/test_quark.py::test_quark_int8_w8a8_moe ## Test Result Before: FAIL…

  • f7ef489 #50960 — [Bugfix] Fix ZMQ port TOCTOU race in shm_broadcast MessageQueue (#50960)
    • 作者: aoshen02 | +68/-4 | 2 个文件

    Why the multi-node topology exposes this The following is the important two-node shape. It has one executor on h200-0 and four GPU workers on h200-1. Each logical queue has one writer; XPUB is the listener and SUB only connects. Before this PR, every XPUB bind marked above used get_open_port() first: it probed a local port, released it, and later bound the real XPUB socket. Port namespaces are …

  • 1c008a3 #51219 — [Bugfix] Close usage telemetry HTTP sessions (#51219)
    • 作者: Nils Matteson | +2/-3 | 1 个文件

    The bug #6600 intentionally moved vLLM’s HTTP callers onto a shared reusable session. That remains useful for repeated asset and media requests. Usage reporting sends during startup and then once every ten minutes. The shared client leaves an external HTTPS socket in the frontend between those reports. ## Measured I tested vLLM d36f24b66 with Qwen3-1.7B on one RTX 4090 Laptop GPU. With usage re…

  • 9823714 #51442 — [Bugfix] Drop stale layer kwarg from online MXFP4 kernel creation (fix precommit) (#51442)
    • 作者: Nick Hill | +0/-1 | 1 个文件

    #49610 removed the layer parameter from make_mxfp4_moe_kernel, but #49347 landed concurrently with a call site still passing it.

  • 70456e5 #50937 — [Bugfix] when loading weights skip empty expert bias if model does not support them (#50937)
    • 作者: Walter Beller-Morales | +89/-1 | 2 个文件

    After the weight loading refactor in https://github.com/vllm-project/vllm/pull/47058 models with unused expert bias tensors started to fail on load with AttributeError: ‘RoutedExperts’ object has no attribute ‘w2_bias’ This is because some quantized exports (particularly those that use GPTQ or llm-compressor) include empty bias even for models that don’t need them. The refactor removed some of the…

  • ebb2972 #51402 — [ROCm][CI][Bugfix] Do not microbatch a step that splits a prefix from its writer (#51402)
    • 作者: stefankoncarevic | +48/-0 | 1 个文件

    tests/v1/distributed/test_dbo.py::test_dbo_dp_ep_gsm8k[deepep_high_throughput] fails in AMD CI roughly every other run. It is not a noisy measurement: the outcome is decided when the server starts and then holds for its life, either around 0.65 on GSM8K or around 0.60, with nothing in between. A degraded server truncates part of its generations after four to six tokens, as if those requests never …

  • 34c1cd2 #49373 — [Bugfix][ROCm] Fix ROCM_AITER_FA & ROCM_AITER_UNIFIED_ATTN QK-Norm+RoPE+KVCache fusion for the packed KV-cache [BLOCKS, HEADS, BLOCK_SIZE, 2*HEAD_DIM] layout (#49373)
    • 作者: Jack Hu | +27/-14 | 3 个文件
    • FA fused do_qk_norm_rope_kvcache_update used stale unbind(1) (broke on the (nb, nkvh, bs, 2*hs) packed layout, num_kv_heads != 2); fixed via a shared _split_kv_cache (transpose(1,2).split) matching unified. - FA fused_qk_norm_rope_kvcache_supported is now gated and not is_shuffle_kv_cache_enabled() — the aiter fused kernel doesn’t support the packed layout in shuffle mode - Corrected kv_stride_o…
  • 448344c #50965 — [Bugfix] Fix get_open_port() livelock on DP-reserved ports and cover get_open_ports_list (#50965)
    • 作者: aoshen02 | +115/-8 | 2 个文件

    Fixes #50024. get_open_port() must return a port that is bindable and outside the range reserved for the data parallel master (VLLM_DP_MASTER_PORT, +10)). The old implementation rejected a reserved candidate and retried without advancing any state. With VLLM_PORT set the scan is deterministic — same start, same answer — so a free-but-reserved candidate (e.g. VLLM_PORT == VLLM_DP_MASTER_POR…

  • 47228db #48534 — [Bugfix][KV-transfer] MoRIIO: per-layer READ-completion barrier in wait_for_layer_load (#48534)
    • 作者: limeward | +86/-13 | 2 个文件

    In MoRIIO READ mode the decode worker posts a request’s per-layer RDMA reads asynchronously in start_load_kv. Each read completes later when a CQ-poll thread flips its status to Succeeded. wait_for_layer_load() was a no-op (pass), so the attention kernel for a layer could run before that layer’s remote KV had actually landed. Under a full CUDA graph the missing barrier is silently skip…

  • d4ecb75 #48758 — [PD][NixlPush][Bugfix] Fix prefix caching (#48758)
    • 作者: Nicolò Lucchesi | +189/-73 | 5 个文件

    Addressing C3 here https://github.com/vllm-project/vllm/issues/48633. On this mode, P needs to send the full sequence to D, regardless of local prefix cache hits; on the other hand, D must take local hits into account, and only request blocks that does not have locally (uncomputed). This means that on P we need to tail-trim and only WRITE/send the last N missing/uncomputed blocks. (side-note: …

  • a231c5c #51260 — [Bugfix] Skip fetching revision for model when model and weights_model are different (#51260)
    • 作者: music-dino | +27/-7 | 2 个文件

    After PR #49990, two gguf tests soft failed on the nightly https://buildkite.com/vllm/ci/builds/82629/list?sid=019fd5a9-db3f-42ba-9b21-993c983b2c7d&tab=output - plugins_tests/gguf/test_gguf_plugin_generate.py::test_models[1-8-32-bfloat16-model0] - plugins_tests/gguf/test_gguf_plugin_generate.py::test_models[1-8-32-bfloat16-model1] with the following error: ### Root cause ModelConfig can load confi…

  • 4f76c8a #50393 — [Bugfix][Platform] Stop re-initializing NVML on every device-capability check (fixes #50381) (#50393)
    • 作者: Sebastian Woo | +82/-1 | 2 个文件

    Fixes #50381. NvmlCudaPlatform.has_device_capability (which CudaPlatform resolves to wherever NVML is available) paid a full nvmlInit()/nvmlShutdown() pair on every call while doing nothing that required an NVML context. ## Root cause The two sibling methods in vllm/platforms/cuda.py were decorated inconsistently: with_nvml_context is just nvmlInit() → call → nvmlShutdown(). The body of has_device…

  • b8db7f4 #50833 — [Bugfix][Quantization] Fix dynamic INT8 W8A8 MoE config being built as W8A16 (#50833)
    • 作者: Hank_ | +5/-4 | 1 个文件

    Fix the runtime quantization config for compressed-tensors INT8 W8A8 MoE models with dynamically quantized per-token activations. TL;DR: For the Triton and CPU INT8 MoE backends, this configuration is currently misclassified as W8A16 while constructing FusedMoEQuantConfig. As a result, the per-token INT8 activation quantization requested by the checkpoint is disabled. For int8_w8a8 compres…

📦 Other

  • 653ebb5 #40116 — Add torch compile for qwen3_vl encoder (#40116)
    • 作者: Tianyu Guo | +22/-1 | 1 个文件

    torch.compile can accelerate the computation of the encoder ## Test Result —

  • 700d39b #50949 — [CPU] Optimize routed FP8/MXFP4 MoE GEMM dispatch (#50949)
    • 作者: Tianmu Li | +110/-5 | 3 个文件

    Remove the fixed global crossover from routed CPU FP8/MXFP4 MoE dispatch so that concentrated moe routing can benefit from brgemm. Select BRGEMM from each expert block’s actual M and N and cover small expert blocks with the existing fixed-row TinyGEMM kernels. Dense and shared-expert dispatch remains unchanged. This is separate from #49942, which adds broader CPU FP8 W8A8 support through different…

  • 7f58e82 #51196 — [Kimi][MM] disable kimi_vit’s dynamic torch.compile for TPU (#51196)
    • 作者: Linkun | +1/-1 | 1 个文件

    Conditionally decorate torch.compile because XLA doesn’t support dynamic compiling. on TPU, the assertion is gone and ViT works without other change. ## Test Result — - [X] The purpose of the PR, such as “Fix some issue (link existing issues this PR will resolve)”.

  • c39076f #51440 — [CI Test] Add specific unit test for mrv2 offloading (#51440)
    • 作者: Wentao Ye | +58/-10 | 2 个文件

    A following up PR for https://github.com/vllm-project/vllm/pull/51413 We add a specific unit test for MRv2 offloading, and revert the non-meaningful ci update from previous PR pytest tests/basic_correctness/test_cpu_offload.py::test_mrv2_weight_offloading With MRv2 change Without CC @njhill

  • c8fe1d5 #51310 — [Spec Decode] Register Qwen3.6 dSpark acceptance coverage (#51310)
    • 作者: Michael Goin | +125/-103 | 2 个文件

    Add dSpark acceptance-rate and GSM8K regression coverage for RedHatAI/Qwen3.6-35B-A3B-speculator.dspark with RedHatAI/Qwen3.6-35B-A3B-NVFP4. The existing per-model tests are registered as data-backed parameterized cases, so future target/draft architectures can supply their model, evaluation, and reference-metric settings without duplicating the test flow. The GSM8K helper also accepts chat-templa…

  • 4a11ed5 #51298 — [DSv32/GLM Perf] Skip short prefill topk for dense mha layer, 97.9% kernel level latency reduction (#51298)
    • 作者: Wentao Ye | +11/-1 | 1 个文件

    We have the logic of skiping short prefill topk in vllm_source/vllm/model_executor/layers/sparse_attn_indexer.py But this is not enabled for GLM5.2 non-torch compiled path as dense_mha_metadata_layer_name="" This PR fixes the issue Acc covered in unit tests Perf can be seen in this ai generated script And we have

  • 99950b0 #51389 — [Profiler] Stamp vLLM version/commit into torch profiler trace metadata (#51389)
    • 作者: Elvir Crnčević | +32/-0 | 1 个文件

    Stamps the vLLM version (and version tuple) into the torch profiler trace metadata via torch.profiler.profile.add_metadata_json, written right after profiler.start() in TorchProfilerWrapper._start. For source installs the setuptools_scm version string embeds the git commit hash, so this effectively records which commit produced the trace. Trace consumers can then pin the file:line source locat…

  • 1c94e8d #45187 — Add NVFP4 KV 4-over-6 scale search (#45187)
    • 作者: Wei-Ming Chen | +284/-42 | 11 个文件

    Add nvfp4_4over6 as an optional KV cache dtype for NVFP4 KV cache stores. The existing nvfp4 behavior is unchanged. Enable 4-over-6 scale search with: For each 16-element NVFP4 group, nvfp4_4over6 evaluates scales derived from max / 6 and max / 4, then stores the candidate with lower reconstruction error. The physical cache layout and capacity remain NVFP4; the existing NVFP4 attention read path i…

  • 3518110 #49347 — [Online quantization] Add online MXFP4 quantization support (#49347)
    • 作者: fxmarty-amd | +1423/-35 | 18 个文件

    This PR adds online MXFP4 linear/MOE quantization support. This lets users pass –quantization mxfp4 (or –quantization online –quantization-config ‘{“linear”: “mxfp4”, “moe”: “mxfp4”}’) on an unquantized (bf16/fp16) model directly. This PR also implements a Triton fallback path (downcast_to_mxfp in mxfp4_utils.py) for MXFP4 dynamic quantization, and a single mxfp4_quantize functions that gathers…

  • a671679 #51413 — [MRv2 Feature] MR v2 weight offloading support (#51413)
    • 作者: Wentao Ye | +19/-3 | 3 个文件

    Part of https://github.com/vllm-project/vllm/issues/41286, reuse current MRv1 offloading, alternative for https://github.com/vllm-project/vllm/pull/51397 Covered in current unit test

  • 56a4b63 #50585 — [K3 Perf] Optimize k3 dspark fused kv, 4.5~4.6x kernel performance improvement (#50585)
    • 作者: Wentao Ye | +104/-88 | 2 个文件

    So we remove the redundant Q calculation and reduce the amount of projections Acc covered in unit tests Perf can be seen in this AI generated script And we get I don’t have enough GPU resources to run an end-to-end benchmark. If you have available GPU capacity and are willing to share access, I would greatly appreciate it!

  • 4eccf90 #44857 — [Attention] Mamba attention module refactor - Final part (#44857)
    • 作者: wangxiyuan | +3/-3 | 1 个文件

    Following #43556. This is the final part for mamba attention module refactor. what’s done in this PR: 1. Change ShortConv from CustomOp to PluggableLayer ## Test Result —

  • 45273b8 #51408 — [1/N] Harden Transformers modelling backend multi-modal path (#51408)
    • 作者: Harry Mellor | +204/-52 | 4 个文件

    Fixed - Text-only prompts crashed multimodal models with KeyError: “Modality ‘image’ not found”. apply() ran _apply_vision whenever the model supported images, not when items were present; the mm_token_type_ids is None guard never fires because HF returns an all-zero tensor. Repro: BAAI/Emu3-Chat-hf + a plain text prompt - Token-ID prompts had special tokens added twice on re-tokenisation: Ge…

  • e229fdb #50007 — [ROCm] Add tuned selective_state_update float32 config for AMD Instinct MI325X (#50007)
    • 作者: vanshbhatia-amd | +51/-0 | 1 个文件

    Adds a tuned selective_state_update (Mamba SSU decode kernel) launch config for the AMD Instinct MI325X for the cache_dtype=float32 variant, mirroring how the MI300X (#47947), MI355 (#47943, #48373) and MI350 (#48159) configs were done. vLLM bundles per-device SSU configs for NVIDIA parts (B200, GB200, H100, H200, RTX PRO 6000) and, more recently, MI300X / MI350 / MI355. On MI325X get_ssm_conf…

  • 7e85d3a #50126 — [ROCm] Enable pinned memory on supported WSL2 kernels (#50126)
    • 作者: Flora Cui | +29/-3 | 2 个文件

    Assisted-by: GitHub Copilot RocmPlatform did not override is_pin_memory_available(), so ROCm builds running under WSL always fell back to the conservative base Platform.is_pin_memory_available(), which unconditionally disables pinned memory. This adds a RocmPlatform override that mirrors the existing kernel-version gating already used by CudaPlatformBase. While touching this code path, the WSL war…

  • f2bfad9 #50068 — [Model] Enable Qwen3.8 for AMD Rocm (#50068)
    • 作者: haic0 | +10/-0 | 1 个文件
    • register the text-only Qwen3_5ForCausalLM and Qwen3_5MoeForCausalLM architectures - advertise hybrid and M-RoPE support on the causal implementation - expose Gated DeltaNet Mamba cache dtype, shape, and copy metadata so text-only Qwen3.5-compatible checkpoints such as Qwen3.8 Max FP8 can initialize through the causal LM path
  • d5aae2b #51357 — Fix ROCm architecture import on non-ROCm platforms (#51357)
    • 作者: Xiaochang Wu | +20/-9 | 2 个文件
    • guard the ROCm-only on_gfx1250 import with current_platform.is_rocm() - cache the architecture check for AITER backend selection - prevent XPU model loading from initializing torch.cuda through the ROCm platform module ## Testing - pre-commit run –from-ref origin/main –to-ref HEAD - DeepSeek V4 loaded all 46 checkpoint shards and completed generation successfully on TP8 XPU
  • ae934ba #48355 — feat: extended EPLB support for Mistral Large 3 and additional MoE backends (#48355)
    • 作者: Julien Debache | +500/-96 | 16 个文件

    Enable EPLB for MoE models whose quantization config derives per-expert state at load time, and for multi-modal models that nest the MoE language model: 1. Nested MoE models. Add get_mixture_of_experts_model() to resolve the MixtureOfExperts interface through VLM wrappers that don’t implement it themselves. The model runner resolves it once and reuses it. 2. **NVFP4 MoE (compressed-tensors W4A…

  • 8d9b52f #51365 — [XPU] quick fix online quantization UT break (#51365)
    • 作者: Yan Ma | +11/-6 | 1 个文件

    Test Result —

  • 5ec47f3 #50234 — [PD][PushConnector] Record last activity of remotes to allow clean up of stale ones (#50234)
    • 作者: Nicolò Lucchesi | +59/-2 | 3 个文件

    Part of the stability enhancement efforts described here https://github.com/vllm-project/vllm/issues/48633. The PushConnector currently inherits all the logic from base_worker.py necessary to clean up old/stale remotes data structures here and prevent cpu leak (more info here https://github.com/vllm-project/vllm/pull/44424). It is not tapping into that workflow though as the connector itself isn’t…

  • da78833 #47972 — Support DeepSeek-V4 AMD Quark NVFP4 with emulation kernel (#47972)
    • 作者: jimmy-adams | +332/-17 | 11 个文件

    Add support for DeepSeek-V4 AMD Quark mixed-quantized checkpoints, where different parts of the model may use different quantization layouts, including per-block FP8 linear layers and NVFP4 MoE experts. This PR focuses on the Quark/DeepSeek-V4 model-side pieces needed to correctly identify and load these mixed-quantized checkpoints. It avoids changing generic weight-loading behavior and keeps Deep…

  • 21ea5b4 #50902 — [rl] Stateful Trainer Send: NCCL + Sparse NCCL [3/N] (#50902)
    • 作者: Aaron Hao | +1418/-779 | 14 个文件

    Trainer-side weight transfer (3/N): migrate NCCL + sparse NCCL ## Context Third PR of the trainer-side weight-transfer rework. - PR 1 (merged, #48042): introduced the new trainer-side abstractions (WeightSource / ModuleSource, VLLMWeightSyncClient, TrainerWeightTransferEngine, WeightTransferTrainerFactory). Purely additive; no backend migrated. - PR 2 (merged, #48981): migrated the **IPC…

🧪 CI/Tests

  • 643c125 #50805 — [ROCm][CI] Baseline legacy extensions in the Torch ABI audit (#50805)
    • 作者: Andreas Karatzas | +7/-1 | 1 个文件
    • Add the exact _C.abi3.so and _rocm_C.abi3.so module names to the temporary unstable-library allowlist. - Keep the allowlist specific to the two remaining ROCm legacy bindings rather than excluding all ROCm extensions. - Retain the stable ABI audit for every extension already migrated to _C_stable_libtorch. PR #48164 enabled the Torch stable ABI audit but did not account for the two ROCm modules …
  • 44351f8 #51410 — [CI] Refresh hybrid Model Runner V2 coverage (#51410)
    • 作者: Jiangyun Zhu | +8/-16 | 5 个文件
    • remove stale Model Runner V2 skips for Qwen3.5 hybrid MTP and Jamba pipeline parallelism - add targeted NVIDIA CI coverage for Qwen3.5 hybrid MTP under Model Runner V2 - fix missing synthetic rejection sampler source dependencies in NVIDIA and Intel CI - replace obsolete torchao nightly/0.12 CI comments with the current pinned-version rationale ## AI assistance AI assistance was used to audit th…
  • 12da9b2 #51451 — [CI] Guard remote-code Transformers compatibility (#51451)
    • 作者: Andreas Karatzas | +7/-0 | 1 个文件
    • Limit Intern-S1-Pro vLLM registry coverage to Transformers 5.14.1. Intern-S1-Pro relies on external trust_remote_code whose video processor imports BASE_VIDEO_PROCESSOR_DOCSTRING, removed in Transformers 5.15. The failure is visible in build 11817’s initialization and processing jobs.
  • 27d7303 #51417 — [CI] Fix Batch Invariance (B200) (#51417)
    • 作者: Jiangyun Zhu | +1/-1 | 1 个文件

    fix https://buildkite.com/vllm/ci/builds/82727#019fdb45-6195-4109-8e19-8a67b7415bce pin quack-kernels==0.6.1 ## Test Result —

  • 0df620d #51288 — [Test] Add packed DeepSeek-V4 KV zeroer geometry regression (#51288)
    • 作者: coltonottley | +195/-0 | 1 个文件

    Adds a CPU-only constructor-crossing regression for the packed DeepSeek-V4 KV zeroer geometry fixed by #50276. The test instantiates the real production seams end to end: - _get_packed_kv_cache_layout - _reshape_attention_kv_cache - DeepseekV4FlashMLABackend - AttentionGroup / KVBlockZeroer.init It verifies that the constructor metadata separates the full packed-row block stride from the meani…

  • a0056e1 #50930 — [Test] Add ROCm AITER MLA op registration and env gating tests (#50930)
    • 作者: Aarushi Jain | +83/-0 | 2 个文件

    [Test] ROCm AITER MLA op registration, fake-tensor, and env gating tests Adds kernel-level tests verifying that rocm_aiter_mla_decode_fwd custom op is correctly registered, supports fake tensors for torch.compile tracing, and respects VLLM_ROCM_USE_AITER / VLLM_ROCM_USE_AITER_MLA environment variable gating. ### Tests added (9 total) - Op registration existence and callable check - mutates_args…

  • 0de0362 #48847 — [ROCm][CI] Loosen block-FP8 fused MoE test tolerance for large-K shapes (#48847)
    • 作者: stefankoncarevic | +56/-11 | 2 个文件

    tests/kernels/moe/test_block_fp8.py::test_w8a8_block_fp8_fused_moe compares two Triton block-FP8 fused-MoE kernels (fused_experts and modular_triton_fused_moe) against a native torch reference. On ROCm/gfx950 (MI355X) the large-K (K=7168), large-N (N >= 1024) DeepSeek-style shapes failed the comparison at the base tolerance. Digging in, the divergence turned out to be **reference artifacts, not ke…

⚡ Performance

  • e644c8c #51434 — [Perf] Optimize DeepSeek V3.2 sequence parallelism (#51434)
    • 作者: Woosuk Kwon | +222/-77 | 3 个文件
    • keep DeepSeek V3.2 hidden and residual states sequence-sharded across every decoder layer, including the dense prefix - use the common optimized sequence-parallel gather/reduce-scatter helpers around attention - replicate the three dense MLPs under sequence parallelism, matching the Kimi K3 and DeepSeek V4 dataflow - shard MTP inputs before eh_proj and restore the full output only once - add foc…
  • a801e71 #48735 — [Perf] Improve --linear-backend filtering (#48735)
    • 作者: Andrii Skliar | +37/-48 | 1 个文件

    The –linear-backend currently raises at startup for any linear layer type the backend doesn’t cover, so single-scheme backends are unusable. flashinfer_b12x is doubly blocked: its only kernel is absent from the NVFP4 auto-selection list, so the filter comes back empty even for NVFP4 layers and the documented opt-in never worked. 1. Replace the five duplicated filter blocks with one _re…

  • 46b5864 #51425 — [Perf] Narrow DeepSeek V3.2 eager CUDA graph region (#51425)
    • 作者: Woosuk Kwon | +61/-16 | 2 个文件

    The portable DeepSeek-V3.2 attention path currently decorates the entire fused attention method with eager_break_during_capture. Only the sparse indexer and backend forward_mqa call require that eager boundary, so query preparation, fused norm/RoPE, and other graph-safe work are replayed as direct launches on every layer. This change: - inlines the graph-safe part of _fused_attention into forward;…

  • 021b7d9 #49390 — [Perf] Raise Blackwell CUDA graph capture default to 1024 (#49390)
    • 作者: Lucas Wilkinson | +48/-8 | 3 个文件

    Raise the default max_cudagraph_capture_size ceiling from 512 to 1024 on data center Blackwell GPUs (compute capability 10.x). Other platforms retain the 512 default, and explicit user configuration remains unchanged. The existing max_num_seqs * decode_query_len * 2 bound still applies, so workloads whose calculated capture size is 512 or smaller are unaffected. On B300, a DeepSeek-V4-Flash DSpark…

🦀 Rust Frontend

  • ac70ce9 #50916 — [Frontend] Disable uvicorn signal handlers instead of racing them (#50916)
    • 作者: Nick Hill | +18/-15 | 3 个文件

    Follow-up to #49668: subclass uvicorn.Server with a no-op capture_signals so uvicorn never installs SIGINT/SIGTERM handlers, instead of polling server.started to order registrations. Removes the startup delay and uvicorn’s handler restoration on exit, and fixes the same override race in DPSupervisor.

🔧 Refactor

  • fcde8e1 #49610 — [Refactor] refactor humming linear and moe backends to use explicit layer configs (#49610)
    • 作者: Jinzhen Lin | +448/-341 | 37 个文件

    This PR refactor the Humming linear and MoE backends to use explicit layer configs and tensors instead of passing vLLM layers into the backend. Also update Humming CI coverage and pin the dependency to the latest upstream commit. AI assistance (OpenAI Codex) was used.

  • c84789c #51051 — [Refactor] Remove kernel dead code (#51051)
    • 作者: Wentao Ye | +0/-396 | 8 个文件

    Remove kernel dead code

🔩 Misc

  • 6b5bec7 #45694 — [Misc] Add and enable Triton kernel unit tests on XPU (#45694)
    • 作者: pmanczak | +58/-13 | 5 个文件

    Makes five Triton kernel unit tests run on Intel GPU (XPU) as well as CUDA, and adds coverage for the round_int8 kernel. Part of RFC #48480. Test-only; no kernel or production code touched. ## Test Result Tested on Arc Pro B70, torch 2.12.0+xpu, triton-xpu 3.7.1 and on H200 on B70: | file | passed | | — | — | | test_block_int8.py | 32 | | test_int8_kernel.py | 44 | | test_triton_scaled_mm.py |…

✨ New Feature

  • 58fcaa0 #49644 — [Feat][Core] Add disk offloading support to SimpleCPUOffloadConnector (#49644)
    • 作者: Guanyi Chen | +466/-19 | 4 个文件

    Implements the disk backend extension envisioned in RFC #19854 (pluggable offloading backends). While #45036 provides SSD offload via the external Mooncake Store connector, this PR adds a native vLLM disk offloading path with zero external dependencies — targeting scenarios where host DRAM is limited but local NVMe capacity is abundant (e.g., dense inference nodes with most memory reserved for…