CalibForge:通过对抗性求解器校准扩展可学习终端任务
原标题:CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
AI 摘要
CalibForge是一个自主终端任务合成系统,通过对抗性求解器校准来调整候选任务,使其处于求解器相对的可学习区间。该系统使用多求解器校准和对比求解器校准两种策略,构建了5,431个校准终端任务。在Terminal-Bench 2.0上,使用完整集合训练的模型达到32.58%和47.57%的准确率,相比基础模型最大提升24.71个百分点,在SWE-bench Pro和Doc2Repo上分别提升27.68和30.04个百分点。研究结果表明,求解器相对的可学习性是构建有效且可迁移的智能体训练数据的实用目标。
正文节选
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks Abstract Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks throu