逐步扩展:大规模混合专家模型的计算高效超参数迁移
原标题:Paper page - Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts
AI 摘要
一篇论文提出了一种计算高效的两步超参数迁移框架,用于预测大规模混合专家(MoE)模型的最优学习率。该方法首先通过宽度缩放迁移学习率,然后外推至万亿级token训练范围,并在155B总参数(17B激活参数)的模型上成功预训练了10万亿token。实验表明,通过小规模代理模型即可高保真地预测大规模训练的最优配置,显著降低消融成本。
正文节选
Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts Abstract A two-step hyperparameter transfer framework predicts optimal learning rates for large Mixture-of-Experts models by scaling across widths and token budgets, enabling efficient pretraining without costly sweeps. Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost. However, optimizing their hyperparameters---par