Gemma Scope 2:开源可解释性工具套件,助力 AI 安全研究
原标题:Gemma Scope 2: helping the AI safety community deepen understanding of complex language model behavior
AI 摘要
Google DeepMind 发布了 Gemma Scope 2,这是一套面向 Gemma 3 全系列模型(270M 至 27B 参数)的开源可解释性工具,旨在帮助 AI 安全社区深入理解复杂语言模型的内部行为。该工具集包含稀疏自编码器和转码器,并采用 Matryoshka 训练技术,支持对越狱、拒绝机制和思维链忠实性等行为进行分析。据称这是 AI 实验室迄今最大规模的开源可解释性工具发布,涉及约 110 PB 数据和超过 1 万亿参数训练。
正文节选
Announcing a new, open suite of tools for language model interpretability Large Language Models (LLMs) are capable of incredible feats of reasoning, yet their internal decision-making processes remain largely opaque. Should a system not behave as expected, a lack of visibility into its internal workings can make it difficult to pinpoint the exact reason for its behaviour. Last year, we advanced the science of interpretability with Gemma Scope, a toolkit designed to help researchers understand th