长时程智能体的跨基准泛化研究
原标题:Cross-Benchmark Generalization in Long-Horizon Agents
AI 摘要
该研究针对长时程智能体的跨基准泛化问题,对开放权重混合专家模型Qwen3.5-122B-A10B进行后训练,使用363个MCP任务和两阶段SFT-RL流程。训练后的模型在五个外部基准上均有提升,包括Toolathlon、tau^2-Bench、BFCL-V4、SWE-Bench Pro和Terminal-Bench 2,其中软件工程基准的提升尤为显著,尽管训练数据不含此类任务。轨迹分析揭示了四种可迁移的行为差异,表明长时程多工具后训练能改变工作方式并跨领域泛化。
正文节选
Computer Science > Software Engineering Title:Cross-Benchmark Generalization in Long-Horizon Agents View PDF HTML (experimental) Abstract:For reinforcement learning (RL) in self-contained environments, a policy can get rewards by exploiting environment-specific regularities (tool schemas, grader parsing, task templates) rather than by acquiring transferable skill, and an in-distribution holdout shares those regularities. We argue that the discriminating question is behavioral, namely