返回全部动态
SPOT:用于在线蒸馏的稀疏探测与结果校准方法
原标题:SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation
AI 摘要
SPOT 是一种用于在线蒸馏(OPD)的新方法,通过稀疏探测和结果校准目标来改进推理性能。它解决了标准 OPD 中教师不确定性分布不均和局部概率无法预测下游成功的问题。实验表明,在多个学生模型和推理基准上,SPOT 相比 OPD 和 EOPD 在 Avg@8/Pass@8 指标上均有显著提升。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation Abstract On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other plausible continuations. Teacher entropy alone does not reveal whether uncertainty is concentrated among a few plausible next tokens or dispersed over a long probability tail, nor whether the student already represents those candidates well.
发布时间:—
抓取时间:2026-08-11 10:29
来源机构:Hugging Face