Harbor Adapters与Harbor-Index:大规模智能体评估基础设施与精选元数据集
原标题:Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
AI 摘要
Harbor Adapters 和 Harbor-Index 项目为大规模智能体评估提供了统一基础设施和精选元数据集。他们开发了适配器,将超过80个基准测试移植到Harbor框架,并对8个模型在54个基准上进行了大规模评估。Harbor-Index包含82个高质量任务,旨在保持挑战性的同时降低成本。所有代码、结果和数据集均已开源。
正文节选
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation Abstract Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and