返回全部动态

DAPD:双锚定策略蒸馏缓解语言模型特权幻觉

原标题:DAPD: Dual-Anchored Policy Distillation

Hugging Face Daily Papers一手来源研究质量 87

AI 摘要

本文提出了一种名为DAPD(双锚定策略蒸馏)的新框架,用于解决语言模型后训练中在线策略自蒸馏(OPSD)存在的特权幻觉问题。DAPD通过双路径锚定和双源锚定两个层面的锚定机制,消除了教师模型与学生模型在推理时的信息不对称。实验表明,DAPD在Qwen3-4B上平均提升2.00分,在4B和32B规模上分别提升2.69和2.78分,显著优于OPSD。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

DAPD: Dual-Anchored Policy Distillation Abstract On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training. It strengthens the teacher with privileged information but can induce a privilege illusion: the student learns privilege-dependent behavior it cannot reproduce from its inference-time context, yet behaves as if the training-time privileged information remained available, ultimately degrading performance. In this paper, we identify information asymmetry b


发布时间:—
抓取时间:2026-08-04 14:13
来源机构:Hugging Face
阅读原文huggingface.co