返回全部动态
SGLang v0.5.12 发布:全面支持 DeepSeek V4 并新增多项优化
原标题:v0.5.12
AI 摘要
SGLang v0.5.12 发布,重点支持 DeepSeek V4 的完整推理路径,包括多种并行策略、硬件适配、预填充-解码分离、HiSparse 和 HiCache 等特性,并新增多个模型支持如 Intern-S2-Preview、MiniCPM-V 4.6 等。此外,还集成了 TokenSpeed MLA 注意力后端,优化了 FP4 低延迟性能,并迁移到 CUDA 13 的 DeepEP。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
# Highlights - **DeepSeek V4 support**: Full inference path for DeepSeek-V4 (#23882), including: Day-0 Features: #23882 - Parallelism: Tensor Parallelism/Expert Parallelism/Context Parallelism/Data Parallel Attention - Hardware: Nvidia B300/B200/H200/H100/GB200/GB300, AMD MI35X - Prefill-Decode Disaggregation - HiSparse for offloading inactive KV cache to CPU memory - Reasoning parser and Tool Call Parser - DeepGemm and FlashMLA kernels for DeepSeek V4, in
发布时间:2026-05-17 02:23
抓取时间:2026-08-02 00:31
来源机构:SGLang