激活预言机学会不阅读:微调预言机中的概念特异性盲区
原标题:When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles
AI 摘要
Hugging Face 每日论文发布了一篇关于激活预言机(Activation Oracles)的研究。研究发现,在受控的禁忌词猜测场景中,针对隐藏特定概念的模型微调得到的激活预言机,反而会选择性丧失恢复该概念的能力,成为概念特异性反阅读器。尽管该概念在预言机内部仍可线性解码,但失败发生在读出路径上,这揭示了行为泄露、表征可解码性和预言机可表述性之间的分离,对可解释性接口的可靠性提出了质疑。
正文节选
When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles Abstract Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They offer a flexible interface for reading hidden information from model states, especially when relevant information is internally represented but absent or incomplete in visible behavior. However, AOs are themselves learned systems: their answers are shaped by