返回全部动态

DeepSeek V4 Pro 与 GPT-5.6 Sol 对比:级联策略以更低成本提升编码准确率

原标题:🚀 DeepSeek V4 Pro 0813 vs. GPT-5.6 Sol on DeepSWE →

Together AI Blog一手来源研究质量 82

AI 摘要

Together AI 在 DeepSWE 基准上对比了 DeepSeek V4 Pro 0813 与 GPT-5.6 Sol,发现 Sol 在单次尝试中准确率更高(72.7% vs 62.8%),但 Pro 成本仅为 Sol 的 1/35,且在多次尝试后覆盖率反超。通过先运行 Pro、失败时升级到 Sol 的级联策略,可解决 83.0% 的任务,每任务成本 3.35 美元,比单独使用 Sol 更便宜且准确率更高。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Don't pick one. Run DeepSeek V4 Pro 0813 first, escalate to GPT-5.6 Sol when the tests fail. That cascade solves 83.0% of DeepSWE tasks at \$3.35 each. Sol alone solves 72.7% at \$8.37. Ten points better, 60% cheaper. - Sol wins the early attempts. 72.7% pass@1 vs 62.8%, and it holds the lead at pass@2 (81.0% vs 78.5%). - Pro wins the last one. 88.5% pass@4 vs 85.8%. Given four tries, the cheap model finds more of the board. - The price gap is 35x. \$0.24 per rollout vs \$8.37. Per \$100 spent,


发布时间:—
抓取时间:2026-08-19 09:44
来源机构:Together AI
阅读原文together.ai