返回全部动态

vLLM 发布 v0.30.0:新增多模型支持与性能优化

原标题:v0.30.0

vLLM Releases一手来源开源质量 90

AI 摘要

vLLM 发布 v0.30.0,包含来自 315 位贡献者的 762 个提交。新增 DeepSeek-V4.1-Flash、GLM-5.3-Flash、K2-Horizon 等模型支持,并引入 Fast Start 权重缓存、Gumbel-max 水印、HiSparse 主机端稀疏解码等特性。同时优化了 Qwen3.8-Flash-Next 与 Kimi K3 性能,扩展大规模服务与量化能力,并包含多项破坏性变更。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

# v0.30.0 ## Highlights This release features 762 commits from 315 contributors (104 new)! * **New models**: DeepSeek-V4.1-Flash (#56214, #56228, #56208) with the whole KV stored in MXFP8 through the FlashMLA V4.1 record on SM100 (#56893), DeepGEMM Mega-mHC (#56962), and async Engram prefetch with Engram DP sharding (#56512); DeepSeek-V4-Flash-Vision-Exp (#54566), also on ROCm (#55107) and with LoRA (#55897); GLM-5.3-Flash (#53906) with EPLB (#55119); K2-Horizon (#55063); Cohere Compass


发布时间:2026-09-22 13:20
抓取时间:2026-09-22 13:31
来源机构:vLLM
阅读原文github.com