返回全部动态

谷歌推出全球首个双盲 AI 评估,防止基准污染

原标题:Piloting the world's first double-blind AI evaluations

Google DeepMind News一手来源研究质量 82

AI 摘要

Google DeepMind 推出了全球首个针对专有前沿 AI 模型的双盲评估,通过与新加坡 AI 安全研究所、OpenMined、AVERI 和 MLCommons 合作,在加密环境中测试 Gemini Flash Lite 模型,以防止基准污染。该技术利用密码学手段确保外部评估问题在测试前不被模型看到,从而提升评估的完整性和可信度。此举旨在增强政策制定者、研究人员和企业对 AI 基准结果的信任。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Building trust in proprietary model benchmarks using cryptographically secure environments Imagine a student is set to take a high-stakes exam. If they accidentally peek at the test questions in advance, achieving a perfect score is influenced by this knowledge, making it a meaningless accomplishment. To truly measure what they know, they must have no visibility of the test questions until it's time to take the exam. That is the exact challenge the industry faces when evaluating advanced AI mode


发布时间:2026-08-27 20:59
抓取时间:2026-09-07 08:45
来源机构:Google DeepMind
阅读原文deepmind.google