返回全部动态

AhaBench:评估智能体能否从先前经验中持续学习的新基准

原标题:AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning

arXiv cs.LG一手来源研究质量 85

AI 摘要

AhaBench 是一个用于评估语言代理在长时间跨度内能否从先前经验中持续学习的新基准,包含 Aha-Puzzle、Aha-Euler 和 Aha-Vending 三个组件。该基准通过初始分数、经验后分数和学习提升三个指标来区分模型的初始能力、利用显式支持的能力以及持久改进的能力。在八模型评测中,Claude Opus 4.6 在经验后总分上领先(64.3),Gemini 3.1 Pro 紧随其后(63.4),但不同模型在不同组件上表现各异。作者发布了完整的基准任务、生成器、验证器和模拟器代码。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

x AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning Abstract Modern language agents are expected to operate over long horizons: they ask follow-up questions, reuse worked examples, handle tool feedback, and adapt to delayed consequences. Most evaluations still reset the agent after a prompt or score only the final state of one trajectory. AhaBench asks a more operational question: when a fixed model receives useful experience, does its later behavi


发布时间:2026-09-09 12:00
抓取时间:2026-09-09 12:35
来源机构:arXiv
阅读原文arxiv.org