谷歌推出全球首个双盲 AI 评估,防止基准污染
原标题:Piloting the world's first double-blind AI evaluations
AI 摘要
Google DeepMind 推出了全球首个针对专有前沿 AI 模型的双盲评估,通过与新加坡 AI 安全研究所、OpenMined、AVERI 和 MLCommons 合作,在加密环境中测试 Gemini Flash Lite 模型,以防止基准污染。该技术利用密码学手段确保外部评估问题在测试前不被模型看到,从而提升评估的完整性和可信度。此举旨在增强政策制定者、研究人员和企业对 AI 基准结果的信任。
正文节选
Building trust in proprietary model benchmarks using cryptographically secure environments Imagine a student is set to take a high-stakes exam. If they accidentally peek at the test questions in advance, achieving a perfect score is influenced by this knowledge, making it a meaningless accomplishment. To truly measure what they know, they must have no visibility of the test questions until it's time to take the exam. That is the exact challenge the industry faces when evaluating advanced AI mode