AI能否评估AI科学家?基于自动多模型评审的自主研究生成系统基准研究
原标题:Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review
AI 摘要
本研究提出了一种使用前沿大语言模型进行自动同行评审的基准测试协议,用于评估AI科学家系统生成的论文质量。研究对Sakana AI v1/v2、CycleResearcher和Data-to-Paper四个框架生成的60篇论文与FARS公司的15篇基准论文进行了比较,发现FARS论文在原创性、严谨性、清晰度和重要性四个维度上显著优于其他系统。Gemini和Claude评审结果高度一致,而GPT-5.4的评审标准有所不同。该研究为AI科学家系统建立了首个定量基准。
正文节选
Computer Science > Artificial Intelligence Title:Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review View PDF HTML (experimental) Abstract:AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and comparing the quality of AI-generated papers remains an open challenge. We propose and implement a rigorous benchmarking protoc