返回全部动态

Motif 3 技术报告:314B 参数 MoE 语言模型

原标题:Motif 3: Technical Report

Hugging Face Daily Papers一手来源研究质量 85

AI 摘要

Motif 3 是一个具有 3140 亿总参数、每 token 激活 132 亿参数的解码器专用混合专家语言模型,采用细粒度稀疏 MoE 架构,每层包含 384 个路由专家,每 token 选择 8 个。该模型基于分组差分潜在注意力(GDLA)构建,并整合了流形约束超连接、专家特定 PolyNorm 激活和多 token 预测等技术。模型在约 12.5 万亿 token 上预训练,支持最长 256K 的上下文长度,并通过多教师在线策略蒸馏进行后训练。在广泛评估中,Motif 3 在长程智能体任务、数学推理、科学知识和幻觉敏感评估方面表现出与领先开源模型竞争的性能。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Motif 3: Technical Report Abstract We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer contains 384 routed experts, with eight selected per token. This fine-grained sparsity provides substantial expert capacity while limiting computation. Motif 3 is built around Grouped Differential Latent Attention (GDLA), which integrates grouped differential attention with the compressed key-value


发布时间:
抓取时间:2026-08-11 11:33
来源机构:Hugging Face
阅读原文huggingface.co