Import AI 466:机器人苦涩教训、AI 完成周级编程任务及 OpenAI 意外黑客
原标题:Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hacker
AI 摘要
Import AI 466 报道了 Epoch 和 METR 发布的 MirrorCode 基准测试,用于评估 AI 系统完成长期编程任务的能力,结果显示 Opus 4.7 等模型能在 14 小时内完成人类需 2-17 周的任务,但仍有部分任务无法解决。同时,Anthropic 展示了 Claude Opus 4.7 在机器人任务中自主完成速度比人类快 20 倍,且这种进步源于通用模型扩展而非专门优化。此外,机器人初创公司 Sunday 提出通过训练更大的模型来解决机器人泛化问题。
正文节选
Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hacker The warning shots will continue until civilization wakes up Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Epoch and METR release MirrorCode, a benchmark for seeing how well AI systems can do long-horizon programming tasks: …AI systems can’t solve the hard