返回全部动态

长时程智能体的跨基准泛化研究

原标题:Cross-Benchmark Generalization in Long-Horizon Agents

arXiv cs.SE一手来源研究质量 82

AI 摘要

该研究针对长时程智能体的跨基准泛化问题,对开放权重混合专家模型Qwen3.5-122B-A10B进行后训练,使用363个MCP任务和两阶段SFT-RL流程。训练后的模型在五个外部基准上均有提升,包括Toolathlon、tau^2-Bench、BFCL-V4、SWE-Bench Pro和Terminal-Bench 2,其中软件工程基准的提升尤为显著,尽管训练数据不含此类任务。轨迹分析揭示了四种可迁移的行为差异,表明长时程多工具后训练能改变工作方式并跨领域泛化。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Computer Science > Software Engineering Title:Cross-Benchmark Generalization in Long-Horizon Agents View PDF HTML (experimental) Abstract:For reinforcement learning (RL) in self-contained environments, a policy can get rewards by exploiting environment-specific regularities (tool schemas, grader parsing, task templates) rather than by acquiring transferable skill, and an in-distribution holdout shares those regularities. We argue that the discriminating question is behavioral, namely


发布时间:2026-08-04 12:00
抓取时间:2026-08-04 14:35
来源机构:arXiv
阅读原文arxiv.org