GPT-6 Astra 基准测试结果矛盾,ARC-AGI-3 效率超人类引关注
原标题:Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward
AI 摘要
OpenAI 的 GPT-6 Astra 在不同基准测试中表现矛盾:Epoch AI 认为其领先,而 Artificial Analysis 认为其与前任持平。在 ARC-AGI-3 测试中,Astra 首次以高于人类的效率完成任务,ARC Prize 负责人 François Chollet 因此将 AGI 预测提前。Astra 在成本效率上表现突出,但部分测试显示其存在短板。
正文节选
Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward OpenAI's GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front, while Artificial Analysis rates it no better than its predecessor. The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the average human for the first time. ARC Prize chief François Chollet calls the progress "2x faster" than he expected and is moving u