返回全部动态
vLLM v0.22.1 发布:新增 Mellum v2 支持与 Zen CPU 优化
原标题:v0.22.1
AI 摘要
vLLM 发布 v0.22.1 补丁版本,包含 8 个提交,来自 6 位贡献者。该版本新增对 JetBrains Mellum v2 模型的支持,并在 AMD Zen CPU 上通过 zentorch 加速量化线性推理,同时修复了多节点 Ray 数据并行服务、DeepSeek-V4 初始化等问题。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
## Highlights This release features 8 commits from 6 contributors (1 new)! v0.22.1 is a patch release on top of v0.22.0 with targeted bug fixes plus a couple of additions: new model support for JetBrains' Mellum v2, zentorch-accelerated quantized linear inference on AMD Zen CPUs, and fixes for multi-node Ray data-parallel serving, DeepSeek-V4 initialization, and a few model-loading regressions. ### Model Support * New model: JetBrains' **Mellum v2**, an open-weights Mixture-of-Experts
发布时间:2026-06-05 18:10
抓取时间:2026-08-02 00:23
来源机构:vLLM