返回全部动态

vLLM v0.22.1 发布:新增 Mellum v2 支持与 Zen CPU 优化

原标题:v0.22.1

vLLM Releases一手来源产品发布质量 79

AI 摘要

vLLM 发布 v0.22.1 补丁版本,包含 8 个提交,来自 6 位贡献者。该版本新增对 JetBrains Mellum v2 模型的支持,并在 AMD Zen CPU 上通过 zentorch 加速量化线性推理,同时修复了多节点 Ray 数据并行服务、DeepSeek-V4 初始化等问题。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

## Highlights This release features 8 commits from 6 contributors (1 new)! v0.22.1 is a patch release on top of v0.22.0 with targeted bug fixes plus a couple of additions: new model support for JetBrains' Mellum v2, zentorch-accelerated quantized linear inference on AMD Zen CPUs, and fixes for multi-node Ray data-parallel serving, DeepSeek-V4 initialization, and a few model-loading regressions. ### Model Support * New model: JetBrains' **Mellum v2**, an open-weights Mixture-of-Experts


发布时间:2026-06-05 18:10
抓取时间:2026-08-02 00:23
来源机构:vLLM
阅读原文github.com