DeepSeek-V4 Flash 与 GPT-5.6 Luna 对比:成本与编码能力
原标题:Model LibraryDeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and CodingWe ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.
AI 摘要
Together AI 对 DeepSeek-V4 Flash 0731 和 GPT-5.6 Luna 在 DeepSWE 基准上进行了 900 次测试。Luna 在 pass@1 上领先 14 个百分点(67.2% vs 53.3%),但 DeepSeek 成本仅为 Luna 的六分之一,每美元解决任务数是 Luna 的 4.8 倍。采用 DeepSeek 优先、失败时升级到 Luna 的级联策略,可达到 78.9% 的解决率,且比单独使用 Luna 便宜 37%。
正文节选
While GPT-5.6 Luna is the stronger engineer on every quality measure, DeepSeek-V4 Flash 0731 is cheap enough that a DeepSeek-first cascade beats Luna alone on both accuracy and cost. - GPT-5.6 Luna leads DeepSWE pass@1 decisively at 67.2% vs 53.3%, a 14 point gap, and holds the lead at every equal attempt count. - DeepSeek-V4 Flash is the cheapest model on the DeepSWE board: \$0.10 per rollout vs \$0.61, delivering 532 solves per \$100 against Luna's 110. - DeepSeek-V4 Flash fails more cleanly,