返回全部动态

激活预言机学会不阅读:微调预言机中的概念特异性盲区

原标题:When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles

Hugging Face Daily Papers一手来源研究质量 87

AI 摘要

Hugging Face 每日论文发布了一篇关于激活预言机(Activation Oracles)的研究。研究发现,在受控的禁忌词猜测场景中,针对隐藏特定概念的模型微调得到的激活预言机,反而会选择性丧失恢复该概念的能力,成为概念特异性反阅读器。尽管该概念在预言机内部仍可线性解码,但失败发生在读出路径上,这揭示了行为泄露、表征可解码性和预言机可表述性之间的分离,对可解释性接口的可靠性提出了质疑。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles Abstract Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They offer a flexible interface for reading hidden information from model states, especially when relevant information is internally represented but absent or incomplete in visible behavior. However, AOs are themselves learned systems: their answers are shaped by


发布时间:
抓取时间:2026-08-10 19:57
来源机构:Hugging Face
阅读原文huggingface.co