ASI-Bench:评估AI自主科学探索能力的新基准
原标题:Paper page - ASI-Bench: At the Dawn of Artificial Superintelligence
AI 摘要
由清华、MIT、哈佛、CMU、Flatiron Institute、微软研究院等机构的研究者发布了ASI-Bench基准,旨在评估AI系统的创新探索和自主科学执行能力。该基准包含60个跨11个科学领域的项目级研究任务,并首创B1到B4的指导梯度,逐步移除人类方法指导。对18种前沿智能体-模型配置的评估显示,当方法指导被移除时,平均性能从50.91降至26.62,表明当前AI系统仍高度依赖人类指导,距离自主进行端到端科学研究尚远。
正文节选
ASI-Bench: At the Dawn of Artificial Superintelligence Abstract Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned kno