Claude Opus 5 发布:性能接近 Fable 5,引发基准测试争议
原标题:[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable)
AI 摘要
Anthropic 于周五发布了 Claude Opus 5 模型,官方基准显示其性能接近 Fable 5,但独立评估认为其实际表现更优。Epoch 的 ECI 指数显示 Opus 5 得分为 159,略低于 Fable 5 的 161,但在软件工程基准上持平。社区反馈强调其在编码和浏览器自动化等实际任务中的出色表现,同时也有对基准测试有效性的质疑。
正文节选
[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable) ain't nobody beats Anthropic at distilling Fable! In a rare Friday release, Opus 5 took the headlines today. Athrough most of its official benchmarks have it technically beating Fable, the official messaging still says it “comes close”. This mostly reflects the difficulty of Evals - today’s AIE track drop - not reflecting “big model smell” that Anthropic obviously knows Fable retains but can’t measure. Fortunately, indepe