返回全部动态

MERIT:成本感知的长期记忆评估揭示工具型代理的记忆效用与风险

原标题:When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

arXiv cs.AI一手来源研究质量 90

AI 摘要

MERIT 基准测试评估了工具使用型 LLM 代理中长期记忆的边际效用,通过 23,440 个带成本核算的片段发现,记忆可将依赖任务的成功率从 0 提升至 0.55-1.00,但嵌入检索在更新事实时表现不稳定,而结构化事实存储和 LLM 摘要等更新时写入的记忆更稳健。研究还表明,记忆实现的选择可导致任务成功率相差 60 个百分点,且完整重放不经济。该基准、测试框架及全部轨迹已开源。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents Abstract Long-term memory for LLM agents is evaluated today by conversational recall benchmarks (LoCoMo, LongMemEval), which measure question answering over dialogue history, not whether remembered facts change what a tool-using agent does. We present MERIT (Memory Evaluation for Realistic Instrumented Tasks), a benchmark and harness that measures the marginal utility of memory for task-executing agents u


发布时间:2026-09-09 12:00
抓取时间:2026-09-09 12:30
来源机构:arXiv
阅读原文arxiv.org