返回全部动态
vLLM v0.27.0 发布:支持 Kimi K3、Qwen3.5,性能大幅提升
原标题:v0.27.0
AI 摘要
vLLM 发布 v0.27.0 版本,包含 561 个提交,来自 242 位贡献者。该版本新增对 Kimi K3、Qwen3.5、K-EXAONE-2.0 等模型的支持,并升级至 PyTorch 2.13.0,深化 FlashAttention 4 集成,优化 DeepSeek-V4 性能,扩展 Model Runner V2 至非生成任务,并引入容错框架和 Rust 前端 gRPC 控制平面。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
# vLLM v0.27.0 Release Notes ## Highlights This release features 561 commits from 242 contributors (64 new)! * **Kimi K3 support** with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656). * **More new mode
发布时间:2026-08-11 05:18
抓取时间:2026-08-11 05:33
来源机构:vLLM