新研究质疑Anthropic和OpenAI关于自主AI研究的说法
原标题:Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach
AI 摘要
普林斯顿大学与英国AI安全研究所联合发布了一项研究,通过“影子评估”方法测试前沿AI模型(如Claude Opus 4.8和GPT-5.6)的自主研究能力。结果显示,这些模型能完成工程任务,但在研究判断、创造性问题解决和资源管理上存在系统性缺陷,生成的两篇论文均被原审稿人拒绝。该研究对Anthropic和OpenAI关于AI可加速自身研究的说法提出了质疑。
正文节选
Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach Anthropic and OpenAI have been touting their models' ability to speed up AI research. A new experiment using unpublished NeurIPS papers tells a different story. Can AI agents conduct AI research on their own? A new paper from Princeton and the UK AI Security Institute puts that claim to the test. Today's frontier models can handle research engineering, but they fail at the parts of the research process that