返回全部动态

SGLang v0.5.17 发布:支持 Kimi K3 与 MiniMax-H3,多项性能优化

原标题:v0.5.17

SGLang Releases一手来源产品发布质量 88

AI 摘要

SGLang 发布 v0.5.17 版本,包含 582 个 PR,来自 194 位贡献者。主要亮点包括对 Kimi K3 和 MiniMax-H3 的 day-0 支持,以及新的 DCP 通信后端、DWDP 预填充策略、会话感知的 Radix Cache 等性能优化。此外,还新增了多个模型支持,并提升了引擎恢复速度和降低了主机开销。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

# Highlights *582 PRs from 194 contributors.* **Kimi K3 day-0 support**: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linear-attention layers interleaved with 24 MLA layers, and a MoonViT3d vision tower, shipping as a native MXFP4 checkpoint. SGLang serves it from day 0 with DCP, DSpark speculative decoding, chunked-prefill PP with TP decode, KDA-aware prefix caching, HiCache L2 over DCP, LoRA on the quantize


发布时间:2026-08-08 08:19
抓取时间:2026-08-08 09:14
来源机构:SGLang
阅读原文github.com