返回全部动态

vLLM v0.25.0 发布:MRv2 默认、新模型与性能优化

原标题:v0.25.0

vLLM Releases一手来源产品发布质量 88

AI 摘要

vLLM 发布 v0.25.0 版本,包含 558 个提交,来自 232 位贡献者。该版本将 Model Runner V2 设为所有稠密模型的默认执行路径,并移除了 PagedAttention 遗留实现。新模型支持包括 LLaVA-OneVision-2、GLM-5、DeepSeek-V3.2 等,并引入了新的流式解析引擎和通用投机解码。性能优化涵盖 GLM-5.2、DeepSeek 和 Blackwell 硬件,显著提升了吞吐量。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

# vLLM v0.25.0 Release Notes ## Highlights This release features 558 commits from 232 contributors (64 new)! * **Model Runner V2 is now the default for all dense models** (#44443). Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compati


发布时间:2026-07-12 04:06
抓取时间:2026-08-02 00:23
来源机构:vLLM
阅读原文github.com