Google DeepMind 用加密双重盲测解决 AI 基准测试信任问题
原标题:AI benchmarks have a trust problem and Google wants to fix it
AI 摘要
Google DeepMind 推出首个针对专有前沿 AI 模型的加密双重盲测,以防止基准测试数据泄露。该测试使用 Google Cloud 的机密计算技术,确保外部测试数据和模型权重均保持私密,消除了传统评估中需要共享测试提示或模型权重的权衡。试点项目与新加坡 AI 安全研究所合作,对 Gemini Flash Lite 模型进行测试,旨在提高 AI 评估的可信度,尤其适用于网络安全等敏感领域。
正文节选
AI benchmarks have a trust problem and Google wants to fix it Google Deepmind wants to use a cryptographic method to stop AI models from seeing test questions in advance. A pilot project with the Singapore AI Safety Institute and other partners runs a double-blind test on a Gemini model. If a test-taker knows the questions ahead of time, even a perfect score is worthless. Google Deepmind uses this image to describe a core problem in evaluating AI models: benchmark contamination. If a model has a