返回全部动态
LMDeploy v0.16.0 发布:支持新模型并优化性能
原标题:v0.16.0
AI 摘要
LMDeploy 发布 v0.16.0 版本,新增对 InternS2 Mobius、GLM-5.2 等模型的支持,并引入 MoE gate v2、CP attention 修复、TurboMind ViT 支持以及 FP8 优化等特性。同时改进了服务端 API 拆分、缓存报告和性能优化,并修复了多个 bug。该版本还升级至 CUDA 13.0,并扩展了 CI 测试覆盖。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
<!-- Release notes generated using configuration in .github/release.yml at main --> ## What's Changed ### 🚀 Features * Support Interns2 mobius by @RunningLeon in https://github.com/InternLM/lmdeploy/pull/4816 * feat: support GLM-5.2 by @CUHKSZzxy in https://github.com/InternLM/lmdeploy/pull/4737 * Intern-S2-Mobius meta-MoE support, MoE gate v2, CP attention fixes by @lzhangzz in https://github.com/InternLM/lmdeploy/pull/4835 * Add TurboMind ViT support for InternVL and Qwen VL models by
发布时间:2026-08-19 12:43
抓取时间:2026-08-31 00:45
来源机构:Shanghai AI Laboratory