返回全部动态
FM-Bench:长期管理决策基准测试揭示智能体行为决定性能
原标题:Paper page - FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
AI 摘要
FM-Bench 是一个用于评估大语言模型智能体在长期决策任务中表现的基准测试,模拟管理足球俱乐部20年的场景。研究发现,管理行为而非模型规模或计算资源决定性能,且所有模型均未能从拒绝的报价中学习市场隐藏价格。该基准提供了首个大规模智能体对抗评估,代码已开源。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents Abstract FM-Bench evaluates long-horizon decision-making of LLM agents managing a football club over 20 years, revealing that managerial behavior rather than scale or token spend drives performance. Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains la
发布时间:—
抓取时间:2026-08-20 10:16
来源机构:Hugging Face