前沿后训练配方回顾:MOPD 模式与 2026 年模型技术解析
原标题:Frontier post-training recipe review with Finbarr Timbers
AI 摘要
Interconnects AI 播客邀请 Ai2 的 Finbarr Timbers 回顾前沿后训练配方,重点介绍了 2026 年出现的多教师在线蒸馏(MOPD)模式,该模式由 MiMo Flash v2 引入,并被 DeepSeek V4 和 Nemotron 3 Ultra 扩展至 10 多个教师模型。节目还梳理了从 InstructGPT 到 2026 年各模型的后训练历史配方,并讨论了 RL 成本高、冲突多以及专家模型可扩展性等 MOPD 兴起的原因。
正文节选
As I’ve been recapping fundamentals of post-training to wrap up my RLHF / Post-training book I knew I needed to get Finbarr Timbers back on the podcast to talk about the state of play. Over the last few months we’ve had many discussions on what we’d need to do to take an Olmo-style recipe to the frontier, supported by Finbarr’s extensive reading of recent model technical reports. To prepare for this, I put together a summary slide deck on the key post-training recipes historically — the path fro