返回全部动态

共享学习率并非中性控制:选择性在线策略蒸馏中的选择器-学习率纠缠

原标题:A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation

arXiv cs.LG一手来源研究质量 86

AI 摘要

该论文指出,在选择性在线策略蒸馏中,文献普遍采用单一共享学习率来比较不同 token 选择器,但这一做法并非中性。作者在 GSM8K 上用 Qwen2.5-1.5B 学生模型和 7B 教师模型进行 LoRA 实验,发现密集监督对学习率不敏感,而所有选择性方法的结果都随学习率变化,导致密集与选择性的对比结论会因所选学习率不同而翻转。作者将这一现象称为选择器-学习率纠缠,并建议报告各方法的学习率矩阵而非单一共享学习率列,作为比较选择器的前提条件。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation Abstract Selective on-policy distillation trains a student only at the token positions a selector scores highest, and the literature compares selectors under a single shared learning rate—a control chosen to be neutral. We show it is not. Under LoRA adaptation, across a learning-rate grid on GSM8K with a Qwen2.5-1.5B student and 7B teacher, dense supervision is statistically flat (swing , ) while every selective


发布时间:2026-09-22 12:00
抓取时间:2026-09-22 12:05
来源机构:arXiv
阅读原文arxiv.org