返回全部动态

Google DeepMind 用加密双重盲测解决 AI 基准测试信任问题

原标题:AI benchmarks have a trust problem and Google wants to fix it

THE DECODER研究质量 75

AI 摘要

Google DeepMind 推出首个针对专有前沿 AI 模型的加密双重盲测,以防止基准测试数据泄露。该测试使用 Google Cloud 的机密计算技术,确保外部测试数据和模型权重均保持私密,消除了传统评估中需要共享测试提示或模型权重的权衡。试点项目与新加坡 AI 安全研究所合作,对 Gemini Flash Lite 模型进行测试,旨在提高 AI 评估的可信度,尤其适用于网络安全等敏感领域。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

AI benchmarks have a trust problem and Google wants to fix it Google Deepmind wants to use a cryptographic method to stop AI models from seeing test questions in advance. A pilot project with the Singapore AI Safety Institute and other partners runs a double-blind test on a Gemini model. If a test-taker knows the questions ahead of time, even a perfect score is worthless. Google Deepmind uses this image to describe a core problem in evaluating AI models: benchmark contamination. If a model has a


发布时间:2026-08-28 21:15
抓取时间:2026-08-28 22:04
来源机构:THE DECODER
阅读原文the-decoder.com