返回全部动态

Together AI 发布 CoderForge-Preview:最大开源编码智能体数据集

原标题:CoderForge-Preview: SOTA open dataset for training efficient coding agents

Together AI Blog一手来源开源质量 84

AI 摘要

Together AI 发布了 CoderForge-Preview,这是目前最大的开源测试验证编码智能体数据集,包含 258,134 条轨迹(其中 155,144 条成功),覆盖 51,201 个任务和 1,655 个仓库。通过在该数据集上微调 Qwen-3 32B 模型,SWE-Bench Verified 性能提升了 23.0%,达到 59.4% 的 pass@1,在 ≤32B 参数的开源数据模型中排名第一。该数据集旨在推动开源 AI 社区在编码智能体训练方面的进展。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

We release CoderForge-Preview - the largest open test-verified coding agent dataset. By leveraging it to fine-tune Qwen-3 32B, we boost SWE-Bench Verified performance 23.0% above the base model reaching 59.4%, ranking #1 among open-data models in the ≤32B parameter range. As coding agents become increasingly capable, the research community faces a critical bottleneck: the lack of large-scale, high-quality open training data. While proprietary models continue to advance, open-weight alternatives


发布时间:—
抓取时间:2026-08-03 01:13
来源机构:Together AI
阅读原文together.ai