MERIT:成本感知的长期记忆评估揭示工具型代理的记忆效用与风险
原标题:When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents
AI 摘要
MERIT 基准测试评估了工具使用型 LLM 代理中长期记忆的边际效用,通过 23,440 个带成本核算的片段发现,记忆可将依赖任务的成功率从 0 提升至 0.55-1.00,但嵌入检索在更新事实时表现不稳定,而结构化事实存储和 LLM 摘要等更新时写入的记忆更稳健。研究还表明,记忆实现的选择可导致任务成功率相差 60 个百分点,且完整重放不经济。该基准、测试框架及全部轨迹已开源。
正文节选
When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents Abstract Long-term memory for LLM agents is evaluated today by conversational recall benchmarks (LoCoMo, LongMemEval), which measure question answering over dialogue history, not whether remembered facts change what a tool-using agent does. We present MERIT (Memory Evaluation for Realistic Instrumented Tasks), a benchmark and harness that measures the marginal utility of memory for task-executing agents u