Kimi K3 与 GPT-5.6 Sol 对比:成本、编码能力与路由策略
原标题:Kimi K3 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
AI 摘要
Together AI 在 DeepSWE 基准上对比了 Kimi K3 与 GPT-5.6 Sol。GPT-5.6 Sol 在单次尝试(pass@1)上领先,但 Kimi K3 在多次尝试(pass@k)上胜出,且成本低 64%。两者失败模式不同,路由结合可覆盖 95.6% 任务,Kimi 优先级联可达到 85.6% 准确率。Kimi K3 为开源模型,可在 Together AI 上部署。
正文节选
GPT-5.6 Sol edges Kimi K3 on single-shot quality, but Kimi wins on pass@k with k > 1 and costs 64% less per completed task. The two models succeed and fail in different ways, which makes routing between them the strongest play on the benchmark. - Kimi K3 vs GPT-5.6 Sol is close on DeepSWE pass@1: Sol leads 72.7% to 68.5%, a 4.2 point gap. - Give the models more attempts and Kimi K3 pulls ahead. It wins pass@2 (82.0 vs 81.0) and pass@4 (89.4% vs 85.8%). - Kimi K3 is far cheaper: \$4.65 per rollou