返回全部动态

SGLang v0.5.12 发布:全面支持 DeepSeek V4 并新增多项优化

原标题:v0.5.12

SGLang Releases一手来源产品发布质量 88

AI 摘要

SGLang v0.5.12 发布,重点支持 DeepSeek V4 的完整推理路径,包括多种并行策略、硬件适配、预填充-解码分离、HiSparse 和 HiCache 等特性,并新增多个模型支持如 Intern-S2-Preview、MiniCPM-V 4.6 等。此外,还集成了 TokenSpeed MLA 注意力后端,优化了 FP4 低延迟性能,并迁移到 CUDA 13 的 DeepEP。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

# Highlights - **DeepSeek V4 support**: Full inference path for DeepSeek-V4 (#23882), including: Day-0 Features: #23882 - Parallelism: Tensor Parallelism/Expert Parallelism/Context Parallelism/Data Parallel Attention - Hardware: Nvidia B300/B200/H200/H100/GB200/GB300, AMD MI35X - Prefill-Decode Disaggregation - HiSparse for offloading inactive KV cache to CPU memory - Reasoning parser and Tool Call Parser - DeepGemm and FlashMLA kernels for DeepSeek V4, in


发布时间:2026-05-17 02:23
抓取时间:2026-08-02 00:31
来源机构:SGLang
阅读原文github.com