返回全部动态

FinPerMA:面向LLM代理的个性化记忆基准测试

原标题:FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

arXiv cs.AI一手来源研究质量 82

AI 摘要

FinPerMA是一个针对LLM代理的个性化记忆基准测试,基于理论信息和事件驱动,用于评估代理在长期金融咨询场景中维护和更新用户模型的能力。该基准包含2994个问题,覆盖276个用户画像,测试结果显示前沿LLM和多种记忆配置均未达到饱和,最高准确率仅约0.47。分析表明,基于摘要的记忆方法在保留事实细节的同时丢失了偏好信号,简单检索方法在事件冲击后表现优于专用记忆系统。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Computer Science > Artificial Intelligence Title:FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents View PDF HTML (experimental) Abstract:Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons. Existing personalized-memory benchmarks primarily test factual retent


发布时间:2026-08-06 12:00
抓取时间:2026-08-06 21:15
来源机构:arXiv
阅读原文arxiv.org