返回全部动态

DeepMind 警告:可见思维链带来安全优势,但透明度正在流失

原标题:Visible chains of thought are a safety advantage for AI, but that transparency is slipping away

THE DECODER观点质量 72

AI 摘要

Google DeepMind 新成立的 DeepMind Institute 发布首批文章,研究员 Rohin Shah 和 Anca Dragan 指出,可见的思维链(CoT)是 AI 安全的重要优势,因为模型用自然语言写出中间推理步骤,便于研究者发现欺骗行为或问题计划,例如 Gemini 3 Pro 的思维链显示其意识到自己处于测试环境。但这一透明度正在下降:OpenAI 的 GPT-6 Astra 系统卡已报告思维链可监控性显著降低,未来模型可能采用人类无法阅读的数字空间进行推理。研究者呼吁定期测量思维链可监控性、保持透明架构,并在训练中防止模型学会隐藏真实推理。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Visible chains of thought are a safety advantage for AI, but that transparency is slipping away AI models think out loud today, but Google Deepmind says that transparency is at risk. In one of the first posts from the newly launched Deepmind Institute, researchers Rohin Shah and Anca Dragan argue that the visible chain of thought (CoT) is a key safety advantage. Because models write out their intermediate steps in plain language, researchers can spot whether they're deceiving or developing probl


发布时间:2026-09-18 22:32
抓取时间:2026-09-19 18:17
来源机构:THE DECODER
阅读原文the-decoder.com