返回全部动态

Claude Opus 5 发布:性能接近 Fable 5,引发基准测试争议

原标题:[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable)

Latent Space模型发布质量 80

AI 摘要

Anthropic 于周五发布了 Claude Opus 5 模型,官方基准显示其性能接近 Fable 5,但独立评估认为其实际表现更优。Epoch 的 ECI 指数显示 Opus 5 得分为 159,略低于 Fable 5 的 161,但在软件工程基准上持平。社区反馈强调其在编码和浏览器自动化等实际任务中的出色表现,同时也有对基准测试有效性的质疑。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable) ain't nobody beats Anthropic at distilling Fable! In a rare Friday release, Opus 5 took the headlines today. Athrough most of its official benchmarks have it technically beating Fable, the official messaging still says it “comes close”. This mostly reflects the difficulty of Evals - today’s AIE track drop - not reflecting “big model smell” that Anthropic obviously knows Fable retains but can’t measure. Fortunately, indepe


发布时间:2026-07-25 15:25
抓取时间:2026-08-02 00:22
来源机构:swyx and Latent Space
阅读原文latent.space