返回全部动态

AgentOPSD:用于智能体强化学习的递归自蒸馏方法

原标题:AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

Hugging Face Daily Papers一手来源研究质量 84

AI 摘要

AgentOPSD 是一种用于智能体强化学习的无评论家递归回合级信用分配方法,通过聚合 token 级教师-学生对数概率差距并递归更新贝叶斯信念状态,将稀疏结果监督转化为回合级信用信号。在 ALFWorld、WebShop 和 Search-QA 上使用 Qwen2.5 模型(3B 和 7B)评估,AgentOPSD 优于 GRPO 和强自蒸馏基线,在 ALFWorld 上使用 Qwen2.5-7B 达到 89.1% 的成功率。消融研究将收益归因于回合级聚合和依赖历史的递归信念更新。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Abstract Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work introduces privileged self-distillation for credit assignment, providing denser supervision, but it remains unclear how such local signals should represent sequential credit. We prop


发布时间:—
抓取时间:2026-08-07 15:19
来源机构:Hugging Face
阅读原文huggingface.co