潜在意图验证:通过深层特征对抗语义伪装攻击
原标题:Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification
AI 摘要
该研究揭示了大型语言模型(LLM)的安全对齐存在表面性缺陷,易受“语义伪装”攻击,即通过将恶意意图包裹在良性叙事中绕过防护。通过分析Phi-3、Qwen2.5和Gemma-2b等小型语言模型的内部激活轨迹,发现存在“意图视界”,即早期层保留可检测的“危害签名”,而后期层则无法区分。基于此,提出了潜在意图验证(LIV)防御方法,在PKU-SafeRLHF数据集上,相比标准防护,零日攻击检测率提升20-50%,且无需重新训练模型。
正文节选
Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification Abstract Safety alignment in Large Language Models (LLMs) is often superficial, relying on refusal mechanisms that trigger only at the final stages of generation without erasing the foundational knowledge of harmful concepts acquired during pretraining. This study demonstrates that this architectural disconnect leaves models vulnerable to Semantic Camouflage—adversarial attacks that wrap harmful intent in benign narra