返回全部动态

LMDeploy v0.15.0 发布:支持 DeepSeek V4 并优化推理性能

原标题:v0.15.0

LMDeploy Releases一手来源产品发布质量 83

AI 摘要

LMDeploy 发布 v0.15.0 版本,新增对 DeepSeek V4 的支持,并引入长上下文和 MTP 前缀缓存命中、推测解码的引导解码、内存分配器与调度器集成等特性。同时优化了 TTFT、流式响应解析、FP8 MoE 等性能,并修复了多项 bug,包括前缀缓存、Triton FP8 通信、多模态工具消息解析等。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

<!-- Release notes generated using configuration in .github/release.yml at main --> ## What's Changed ### 🚀 Features * Support long-context and MTP prefix-cache hits by @grimoire in https://github.com/InternLM/lmdeploy/pull/4688 * [Feature] Add guided decoding support for speculative decoding by @windreamer in https://github.com/InternLM/lmdeploy/pull/4559 * feat(turbomind): memory allocator, object cache, and scheduler integration by @lzhangzz in https://github.com/InternLM/lmdeploy/pull


发布时间:2026-07-31 21:00
抓取时间:2026-08-31 00:45
来源机构:Shanghai AI Laboratory
阅读原文github.com