FinPerMA:面向LLM代理的个性化记忆基准测试
原标题:FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents
AI 摘要
FinPerMA是一个针对LLM代理的个性化记忆基准测试,基于理论信息和事件驱动,用于评估代理在长期金融咨询场景中维护和更新用户模型的能力。该基准包含2994个问题,覆盖276个用户画像,测试结果显示前沿LLM和多种记忆配置均未达到饱和,最高准确率仅约0.47。分析表明,基于摘要的记忆方法在保留事实细节的同时丢失了偏好信号,简单检索方法在事件冲击后表现优于专用记忆系统。
正文节选
Computer Science > Artificial Intelligence Title:FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents View PDF HTML (experimental) Abstract:Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons. Existing personalized-memory benchmarks primarily test factual retent