共享学习率并非中性控制:选择性在线策略蒸馏中的选择器-学习率纠缠
原标题:A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation
AI 摘要
该论文指出,在选择性在线策略蒸馏中,文献普遍采用单一共享学习率来比较不同 token 选择器,但这一做法并非中性。作者在 GSM8K 上用 Qwen2.5-1.5B 学生模型和 7B 教师模型进行 LoRA 实验,发现密集监督对学习率不敏感,而所有选择性方法的结果都随学习率变化,导致密集与选择性的对比结论会因所选学习率不同而翻转。作者将这一现象称为选择器-学习率纠缠,并建议报告各方法的学习率矩阵而非单一共享学习率列,作为比较选择器的前提条件。
正文节选
A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation Abstract Selective on-policy distillation trains a student only at the token positions a selector scores highest, and the literature compares selectors under a single shared learning rate—a control chosen to be neutral. We show it is not. Under LoRA adaptation, across a learning-rate grid on GSM8K with a Qwen2.5-1.5B student and 7B teacher, dense supervision is statistically flat (swing , ) while every selective