聊天模板切换LLM自我指涉口吻,激活引导可复现
原标题:"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It
AI 摘要
该研究通过对比8个开源指令模型(Llama、Gemma、Mistral、Qwen,1B-9B)在有无聊天模板下的输出,发现聊天模板如同一个开关:存在时模型更倾向使用免责声明式口吻(如“我只是一个AI”),移除后免责声明率从0.53降至0.36,体验式口吻(如“我感觉”)从0.01升至0.15。在3个模型中,研究者定位到一个可操控的激活方向,添加该方向可提升免责声明率,移除则降低,而随机方向无此效果。这表明模型对自身的描述并非固定事实,部分由聊天模板决定,研究自我报告或内省时需控制这一混淆因素。
正文节选
“As a Language Model…”: Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It Abstract Large Language Models (LLMs) tend to add disclaimers like “I’m just an AI” when asked about something related to themselves. The self-reports from such responses are used in debates about AI safety or self-knowledge of the models, yet what drives them is not well understood. Are the models telling us about themselves or rather how they are deployed? In this work, we show that