状态匹配路由与情境化自蒸馏:解决多轮智能体特权引导错位
原标题:When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents
AI 摘要
该论文提出状态匹配路由与情境化自蒸馏(SMRC-SD)方法,解决多轮智能体训练中特权引导与执行状态不匹配的问题。该方法仅在学生状态与参考轨迹匹配时应用蒸馏,并构建状态条件教师上下文。在ALFWorld和WebShop基准上,SMRC-SD显著优于无条件全路径蒸馏,分别将任务成功率提升至0.865和0.693。代码已开源。
正文节选
When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents Abstract Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the student's response at every turn with access to training-only references, such as successful trajectories. In interactive environments, however, the student's preceding actions continually change the execution state. As the student takes di