返回全部动态

逐步扩展:大规模混合专家模型的计算高效超参数迁移

原标题:Paper page - Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

Hugging Face Daily Papers一手来源研究质量 85

AI 摘要

一篇论文提出了一种计算高效的两步超参数迁移框架,用于预测大规模混合专家(MoE)模型的最优学习率。该方法首先通过宽度缩放迁移学习率,然后外推至万亿级token训练范围,并在155B总参数(17B激活参数)的模型上成功预训练了10万亿token。实验表明,通过小规模代理模型即可高保真地预测大规模训练的最优配置,显著降低消融成本。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts Abstract A two-step hyperparameter transfer framework predicts optimal learning rates for large Mixture-of-Experts models by scaling across widths and token budgets, enabling efficient pretraining without costly sweeps. Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost. However, optimizing their hyperparameters---par


发布时间:—
抓取时间:2026-08-24 13:24
来源机构:Hugging Face
阅读原文huggingface.co