返回全部动态
在线策略蒸馏真的有效吗?从噪声教师到自我改进
原标题:Paper page - Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
AI 摘要
该论文研究了在线策略蒸馏(OPD)在推理任务中的实际作用,发现其提升主要来自抑制低概率token而非教师模型的指导。基于此,作者提出了一种无需监督的熵自适应方法OPSA,在AIME24基准上相比基础模型Qwen3-1.7B将Avg@32提升了35.41分,相对提升263%,并显著优于传统OPD方法。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Abstract On-policy distillation relies mainly on suppressing low-probability tokens rather than teacher guidance, motivating a supervision-free entropy-adaptive method that substantially improves reasoning performance. On-policy distillation (OPD) offers dense token-level supervision as an alternative to the sparse outcome-level advantages of reinforcement learning with verifiable rewards (RLVR). However, the teac
发布时间:—
抓取时间:2026-09-01 11:51
来源机构:Hugging Face