返回全部动态
电路引导权重缩放:提升LLM安全性的新方法
原标题:From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling
AI 摘要
该研究从机制可解释性角度揭示了LLM安全行为的电路级组织,包含有害检测头、安全神经元和拒绝头三个阶段。通过跨六种LLM的权重缩放干预,在攻击下安全率提升26.5%,同时标准基准准确率仅下降1.7%。研究提供了因果证据,表明安全行为由跨层组件协同实现,且该结构可迁移。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling Abstract Despite extensive alignment efforts, Large Language Models (LLMs) remain vulnerable to generating unsafe content under adversarial prompting, yet the internal mechanisms by which safety behaviors are implemented remain poorly understood. We study LLM safety from a mechanistic interpretability perspective and characterize a multi-stage safety circuit that organizes refusal behavior, consisting of (i) Harmful Detectio
发布时间:2026-09-02 12:00
抓取时间:2026-09-02 12:10
来源机构:arXiv