返回全部动态
LMDeploy v0.15.0 发布:支持 DeepSeek V4 并优化推理性能
原标题:v0.15.0
AI 摘要
LMDeploy 发布 v0.15.0 版本,新增对 DeepSeek V4 的支持,并引入长上下文和 MTP 前缀缓存命中、推测解码的引导解码、内存分配器与调度器集成等特性。同时优化了 TTFT、流式响应解析、FP8 MoE 等性能,并修复了多项 bug,包括前缀缓存、Triton FP8 通信、多模态工具消息解析等。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
<!-- Release notes generated using configuration in .github/release.yml at main --> ## What's Changed ### 🚀 Features * Support long-context and MTP prefix-cache hits by @grimoire in https://github.com/InternLM/lmdeploy/pull/4688 * [Feature] Add guided decoding support for speculative decoding by @windreamer in https://github.com/InternLM/lmdeploy/pull/4559 * feat(turbomind): memory allocator, object cache, and scheduler integration by @lzhangzz in https://github.com/InternLM/lmdeploy/pull
发布时间:2026-07-31 21:00
抓取时间:2026-08-31 00:45
来源机构:Shanghai AI Laboratory