返回全部动态

SAF-OPD:稳定优势融合用于在线蒸馏

原标题:SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

Hugging Face Daily Papers一手来源研究质量 84

AI 摘要

SAF-OPD 提出了一种稳定优势融合框架,用于结合强化学习(RLVR)和在线蒸馏(OPD)训练。该框架通过四阶段流程(稀疏化、压缩、预热、退火)解决固定系数融合导致的熵崩溃问题,在七个数学推理和代码生成基准上,使用 Qwen3-1.7B/4B/8B 模型,将聚合分数提升了 0.51-2.70%,并实现了更稳定的训练。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation Abstract Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stronger teacher for a dense advantage but caps performance at teacher quality and discourages exploration beyond it. Their complementarity makes combining RLVR and OPD promising, but we find that fusing the two advantages with a fixed coefficient triggers


发布时间:
抓取时间:2026-08-03 20:01
来源机构:Hugging Face
阅读原文huggingface.co