返回全部动态

Apple研究揭示数据受限下混合预训练的扩展规律

原标题:Scaling Laws for Mixture Pretraining Under Data Constraints

Apple Machine Learning Research一手来源研究质量 87

AI 摘要

Apple机器学习研究团队通过超过2000次语言模型训练实验,研究了数据受限条件下混合预训练的扩展规律。研究发现,稀缺目标语料可重复使用15-20次,且混合训练比单一来源训练能容忍更高的重复率。团队提出了一个考虑重复感知的混合扩展定律,可优化计算有效的数据混合配置,为数据受限条件下的预训练提供实用建议。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

As language models scale, the amount of data they require grows – yet many target data sources, such as low-resource languages or specialized domains, are inherently limited in size. A common strategy is to mix this scarce but valuable target data with abundant generic data, which presents a fundamental trade-off: too little target data in the mixture underexposes the model to the target domain, while too much target data repeats the same examples excessively, yielding diminishing returns and ev


发布时间:2026-08-20 08:00
抓取时间:2026-08-20 23:34
来源机构:Apple
阅读原文machinelearning.apple.com