-
-
Notifications
You must be signed in to change notification settings - Fork 22.3k
Pull requests: vllm-project/vllm
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[Bugfix][Frontend] Expand the assistant tokens mask through multimodal placeholder expansion
bug
Something isn't working
#57254
opened Sep 16, 2026 by
he-yufeng
Contributor
Loading…
[Bugfix][Core] Keep cache-registered Mamba states out of retirement
bug
Something isn't working
kv-cache-manager
#57253
opened Sep 16, 2026 by
sungsooha
Contributor
Loading…
[Bugfix][ROCm] Add record_logical_topk_ready to ROCMAiterMLASparseImpl (GLM-5.3-Flash boot crash)
bug
Something isn't working
glm
rocm
Related to AMD ROCm
#57252
opened Sep 16, 2026 by
mustafayildirim
Loading…
[Metrics][KV Offload] Add Prometheus metrics for SimpleCPUOffloadConnector
documentation
Improvements or additions to documentation
kv-connector
[Core] structured generation mode for DiffusionGemma model (Jev-like)
documentation
Improvements or additions to documentation
#57250
opened Sep 16, 2026 by
mmastrac
Loading…
3 of 4 tasks
[ROCm][CI] Fix Nixl+Offloading PD edge cases on TheRock image
ci/build
kv-connector
rocm
Related to AMD ROCm
#57249
opened Sep 16, 2026 by
mawong-amd
Contributor
Loading…
4 tasks
[Bugfix][Parser] Preserve Gemma4 content when reasoning parsing is disabled
bug
Something isn't working
tool-calling
#57247
opened Sep 16, 2026 by
Prudhvivuda
Contributor
Loading…
2
[Bugfix][ROCm][MoE] Fix AITER MoE routing fallback to eager torch.topk for scoring_func=softmax
bug
Something isn't working
rocm
Related to AMD ROCm
#57246
opened Sep 16, 2026 by
shantipriya-amd
Contributor
•
Draft
3 of 4 tasks
[LoRA] Add DeepSeek V4.1 text and vision adapter support
deepseek
Related to DeepSeek models
DSv4
DSv4.1
Related to DeepSeek-V4.1 models
#57245
opened Sep 16, 2026 by
HollowMan6
Contributor
Loading…
3 of 4 tasks
[Bugfix][Frontend] Handle malformed Anthropic tool arguments
bug
Something isn't working
frontend
#57244
opened Sep 16, 2026 by
subhashpolisetti
Contributor
Loading…
[Bugfix][Frontend] Support scalar priority and wire priority into LLM.chat()
bug
Something isn't working
frontend
#57243
opened Sep 16, 2026 by
subhashpolisetti
Contributor
Loading…
[Bugfix] Default missing detail for Responses API input images
bug
Something isn't working
frontend
tool-calling
#57241
opened Sep 16, 2026 by
yzong-rh
Contributor
Loading…
4 tasks done
[WIP][Model][ROCm][DCP] DeepSeek-V4 decode context parallelism
deepseek
Related to DeepSeek models
DSv4
needs-rebase
rocm
Related to AMD ROCm
#57239
opened Sep 16, 2026 by
edwinlim0919
Contributor
•
Draft
3 of 9 tasks
[CI] Shard attention by whole units with six primary and two mirror replicas
ci/build
cpu
Related to CPU backends
deepseek
Related to DeepSeek models
dflash
glm
nvidia
rocm
Related to AMD ROCm
#57238
opened Sep 16, 2026 by
Thangnguyenvn98
Contributor
Loading…
[CI] Folder-shard Spec Decode Speculators + MTP 1→4
ci/build
deepseek
Related to DeepSeek models
qwen
Related to Qwen models
speculative-decoding
#57237
opened Sep 16, 2026 by
Thangnguyenvn98
Contributor
Loading…
[A2A] Allow EPLB + SharedExpert Overlap + DeepEPv2
ready
ONLY add when PR is ready to merge/full CI is needed
#57236
opened Sep 16, 2026 by
robertgshaw2-redhat
Collaborator
Loading…
[Perf][TP] Skip full-vocabulary gather for exact greedy sampling
qwen
Related to Qwen models
#57235
opened Sep 16, 2026 by
WildWestOat
Loading…
[Bugfix][Multimodal] Tolerate malformed EXIF in Molmo 2 image preprocessing
bug
Something isn't working
multi-modality
Related to multi-modality (#4194)
#57234
opened Sep 16, 2026 by
Hotragn
Contributor
Loading…
[Rust Frontend] Bump vllm-proto to 0.3.0
rust
#57233
opened Sep 16, 2026 by
alec-flowers
Contributor
Loading…
[ROCm][Bugfix] Reduce CUDA graph divergences
bug
Something isn't working
nvidia
rocm
Related to AMD ROCm
#57229
opened Sep 16, 2026 by
mawong-amd
Contributor
Loading…
4 tasks
[Core] Don't issue a blocking collective RPC during engine handshake
#57226
opened Sep 16, 2026 by
okorzh-amd
Contributor
Loading…
4 tasks
Cap multimodal content-part uuid length (cache_salt parity)
frontend
#57221
opened Sep 16, 2026 by
AUTHENSOR
Loading…
[Perf][DSV4.1] Prefer MegaMoE for supported local expert parallelism
DSv4.1
Related to DeepSeek-V4.1 models
#57220
opened Sep 16, 2026 by
WoosukKwon
Collaborator
Loading…
Previous Next
ProTip!
Add no:assignee to see everything that’s not assigned.