返回全部动态

GPT-6 Astra 基准测试结果矛盾,ARC-AGI-3 效率超人类引关注

原标题:Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

THE DECODER模型发布质量 76

AI 摘要

OpenAI 的 GPT-6 Astra 在不同基准测试中表现矛盾:Epoch AI 认为其领先,而 Artificial Analysis 认为其与前任持平。在 ARC-AGI-3 测试中,Astra 首次以高于人类的效率完成任务,ARC Prize 负责人 François Chollet 因此将 AGI 预测提前。Astra 在成本效率上表现突出,但部分测试显示其存在短板。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward OpenAI's GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front, while Artificial Analysis rates it no better than its predecessor. The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the average human for the first time. ARC Prize chief François Chollet calls the progress "2x faster" than he expected and is moving u


发布时间:2026-09-04 19:07
抓取时间:2026-09-04 19:39
来源机构:THE DECODER
阅读原文the-decoder.com