RoMeRL:通过降阶效用状态平衡自进化智能体记忆中的反馈覆盖与记忆-奖励陷阱
原标题:RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States
AI 摘要
RoMeRL 提出了一种用于自进化 LLM 智能体的降阶记忆强化学习方法,通过固定维度的语义坐标表示记忆状态,解决反馈分散和记忆-奖励陷阱问题。在 ALFWorld 和 LifelongAgentBench 基准上,该方法将 Cold-Q 比率降低 80.0%,反馈密度提高约 6 倍,内存占用减少 84.4%,LLM 调用减少 21.1%。代码已开源。
正文节选
RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States Abstract Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-indexed utilities grow with the interaction history, thereby dispersing limited feedback over an ever-expanding state space. Second, because trajectory-level rewards are jointly assigned to co-retrieved memories, irrelevant experiences may receive mis