面向软件工程智能体的类别感知迭代专家训练
原标题:Paper page - One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents
AI 摘要
该论文提出面向软件工程智能体的类别感知迭代专家训练框架,通过 Refresh-Repair-Expand 迭代开发各类别专家,并用标签路由的多教师在线策略蒸馏(MOPD)将其整合为单一可部署模型。最终策略在 Pro-618 上达到 58.04%、在 SWE-bench Multilingual 上达到 59.00%,分别较基座模型提升 5.39 和 2.78 个百分点,且无需外部模型提供解题轨迹。模型 Logics-SWE-Qwen3.6-27B 与数据集 Logics-SWE-Env-2.5K 已在 Hugging Face 公开。
正文节选
One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents Abstract Repository-level software engineering (SWE) comprises heterogeneous task categories, whose progress under pooled agentic reinforcement learning can be uneven: gains in some categories coincide with regressions in others, while aggregate resolution obscures these changes. Motivated by this category see-saw, we develop a category-aware expert-training and policy-integration framework. Execut