返回全部动态

inclusionAI 发布 Ling-3.0-flash:124B 总参混合推理模型,激活仅 5.1B

原标题:inclusionAI/Ling-3.0-flash-fp8

Hugging Face New and Trending Models一手来源模型发布质量 88

AI 摘要

inclusionAI 发布了 Ling-3.0-flash,一款原生混合推理模型,总参数 124B,激活参数仅 5.1B,采用混合线性注意力架构(KDA 与 MLA 交替)和稀疏 MoE。该模型在代码、智能体、通用知识等基准上匹配或超越前代旗舰,并集成 SGLang HiCache 与 Mooncake 分层缓存,将长输入场景的 TTFT 降低 60% 至 80% 以上。模型支持 256K 上下文,默认启用思考模式,适用于生产环境中的复杂智能体工作流。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

--- license: mit pipeline_tag: text-generation --- <p align="center"> <img src="https://mdn.alipayobjects.com/huamei_qa8qxu/afts/img/A*4QxcQrBlTiAAAAAAQXAAAAgAemJ7AQ/original" width="100"/> </p> <p align="center">🤗 <a href="https://huggingface.co/inclusionAI">Hugging Face</a>&nbsp;&nbsp; | &nbsp;&nbsp;🤖 <a href="https://modelscope.cn/organization/inclusionAI">ModelScope </a>&nbsp;&nbsp; | &nbsp;&nbsp;🐙 <a href="https://openrouter.ai/inclusionai/ling-3.0-flash:free">OpenRouter </a>&nbsp;&nbsp


发布时间:2026-09-04 13:42
抓取时间:2026-09-04 13:43
来源机构:Hugging Face
阅读原文huggingface.co