AgentOPSD:用于智能体强化学习的递归自蒸馏方法
原标题:AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
AI 摘要
AgentOPSD 是一种用于智能体强化学习的无评论家递归回合级信用分配方法,通过聚合 token 级教师-学生对数概率差距并递归更新贝叶斯信念状态,将稀疏结果监督转化为回合级信用信号。在 ALFWorld、WebShop 和 Search-QA 上使用 Qwen2.5 模型(3B 和 7B)评估,AgentOPSD 优于 GRPO 和强自蒸馏基线,在 ALFWorld 上使用 Qwen2.5-7B 达到 89.1% 的成功率。消融研究将收益归因于回合级聚合和依赖历史的递归信念更新。
正文节选
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Abstract Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work introduces privileged self-distillation for credit assignment, providing denser supervision, but it remains unclear how such local signals should represent sequential credit. We prop