返回全部动态
vLLM v0.28.0 发布:Kimi-K3 与 DeepSeek V4 性能优化
原标题:v0.28.0
AI 摘要
vLLM 发布 v0.28.0 版本,包含 584 个提交,重点优化了 Kimi-K3 和 DeepSeek V4 的性能,支持稀疏 MLA、DCP、FlashKDA 等特性,并改进了推测解码和模型运行器。该版本还引入了新的默认配置、破坏性变更,并扩展了模型支持,如 Muse Glimmer、Ling 3.0 Flash 等。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
# v0.28.0 ## Highlights This release features 584 commits from 270 contributors (76 new)! * **Kimi-K3 performance push**: a major optimization effort for Kimi-K3 across the stack — Decode Context Parallel (DCP) support (#50484), fused FlashKDA decode and prefill kernels (#50654, #51311, #52458), SiTU activation support for MegaMoE (#50510), GEMM-RS for sequence parallelism (#52079), combined all-gathers with 1.5~3x kernel-level speedup (#51070), an adaptive speculative token budget deli
发布时间:2026-08-26 17:46
抓取时间:2026-08-28 10:34
来源机构:vLLM