返回全部动态

RoMeRL:通过降阶效用状态平衡自进化智能体记忆中的反馈覆盖与记忆-奖励陷阱

原标题:RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States

Hugging Face Daily Papers一手来源研究质量 84

AI 摘要

RoMeRL 提出了一种用于自进化 LLM 智能体的降阶记忆强化学习方法,通过固定维度的语义坐标表示记忆状态,解决反馈分散和记忆-奖励陷阱问题。在 ALFWorld 和 LifelongAgentBench 基准上,该方法将 Cold-Q 比率降低 80.0%,反馈密度提高约 6 倍,内存占用减少 84.4%,LLM 调用减少 21.1%。代码已开源。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States Abstract Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-indexed utilities grow with the interaction history, thereby dispersing limited feedback over an ever-expanding state space. Second, because trajectory-level rewards are jointly assigned to co-retrieved memories, irrelevant experiences may receive mis


发布时间:
抓取时间:2026-08-11 11:33
来源机构:Hugging Face
阅读原文huggingface.co